The Standup Problem: Why You Freeze When It's Your Turn
· 6 min read · AptSpeak Team

The person before you says "yeah, that's me." Their tile stops glowing. There's a half-second of nothing, and then someone says your name, and eleven video squares turn toward you at once. You know exactly what you did yesterday. You spent nine hours on it. And for two seconds you produce nothing, and those two seconds stretch out into something that feels like twenty, and when words finally arrive they arrive too fast and in the wrong order.
Then it's over, someone else is talking, and you spend the next four minutes replaying it instead of listening.
You've probably never mentioned this to anyone. It sounds like a small thing to struggle with when you can read a 400-line diff, write a design doc, and argue about database indexes in Slack all afternoon. That's exactly why it doesn't get fixed. It gets filed under "I should be more confident" and left there.
It isn't a confidence problem. It's a format problem, and formats can be changed.
What your brain is actually doing in those two seconds
Speaking in standup asks your working memory to run three processes at the same time.
The first is retrieval. You have to reconstruct yesterday out of a mess of branches, tabs, interruptions, and a meeting that ate an hour. Nobody's memory serves that up cleanly.
The second is production in your second language. You're fluent, so most of this runs cheaply, but "cheaply" is not "free." Choosing tense, choosing between fixed and have fixed, choosing how to compress a messy technical situation into a clause, all of it costs something.
The third is self-monitoring. You're listening to yourself as you speak, checking your accent, checking whether that sentence landed, watching your manager's face in a small rectangle for signs that you're taking too long.
In your first language, two of those three run on autopilot. Production is automatic and monitoring barely happens, which leaves nearly all of your working memory free for retrieval. In English, production costs real resources and monitoring costs more, because there's an audience and a status element. Working memory is a fixed budget. When two processes take more, the third gets less, and retrieval is the one that fails first. That's the blank. It's a load effect, well documented in second-language research, and it has nothing to do with whether you know the words.
Which tells you where the fix goes. You can't make monitoring cheaper by wanting it to be cheaper. You can move retrieval out of the live moment entirely.
The four shapes this takes
Before the fix, it helps to recognise your own pattern. There are four common ones and most people have a favourite.
The info-dump. You list everything. Every file, every rabbit hole, every small discovery, in chronological order, because chronological order is the only structure available when you're improvising. It runs ninety seconds and nobody retains any of it.
The apology opener. "Sorry, um, so yesterday I was just..." You've spent your first four words lowering expectations before you've said anything. It buys you a moment of thinking time, which is why you do it, and it costs you the frame.
The trail-off. You get through most of it and then you can't find the exit, so you land on "...yeah, so, that's it." The update was fine. The ending makes it sound like it wasn't.
The false blocker. Someone asks if you're blocked and you say no, because you are blocked, but explaining the blocker properly would take four sentences and a bit of context, and you don't have four sentences ready. This is the expensive one. It's the only failure here that actually costs the team something.
Three beats: what moved, what's next, what I need
Standup does not need a narrative. It needs three pieces of information, in this order.
Key takeaway
What's next. One thing you're doing today.
What I need. A person, a decision, or an unblock. "Nothing" is a legitimate answer when it's true.
That's it. Thirty seconds is about seventy to ninety words of speech, which is far less than you think.
Here's the same day, twice.
Improvised: "Yeah, sorry, um, so yesterday I was working on the payment thing, I was looking at
checkout_service.pyand also the validation, and there was something weird with the webhook handler, I changed the retry logic and then the tests started failing so I was looking at that for a while, and I also reviewed two PRs, and today I think I continue with that. Um. Yeah. No blockers."
Prepared: "Payments: the retry logic is done and merged. Today I'm fixing four webhook tests that broke with it. I need ten minutes with Marta after this to confirm what retry count staging expects, otherwise I'm guessing."
Same work. Same English. The second one took eight seconds less and told the team something they can act on. Note that the blocker is now in the room, and it cost one sentence.
The thirty-second pre-write
Open a note before the call. Write three lines. What moved, what's next, what I need.
You are not writing a script to read aloud. Reading aloud sounds like reading aloud, and you'll drift from it the moment someone asks a follow-up. You're writing a scaffold. The point is that retrieval happens now, in a quiet moment, with all of your working memory available for it, so that when your name comes it isn't a process anymore. It's a lookup.
The objection is that this feels like cheating, or like overkill for a two-minute meeting. It isn't either. Your product manager writes down their talking points. Your tech lead has a list open during planning. Nobody calls that cheating, because nobody expects a live spoken summary to be improvised from nothing. The improvised version is not more honest. It's just noisier, and it puts the cost on the eleven people listening rather than on you for thirty seconds beforehand.
Do it for two weeks. The pre-write shrinks on its own, because the structure starts loading itself.
When it happens anyway
Some days you'll blank regardless. Tired, context-switched, dragged into the call from something else.
Tip
Then look at your three lines and start.
And know this: a two-second pause is completely invisible from the outside. You experience it from inside your own head, where it's loud and slow and humiliating. Everyone else experiences it as a normal beat between two speakers on a laggy video call. They are not evaluating it. Half of them are reading Slack.
Tomorrow morning
The person before you finishes. Your name comes. This time the work of remembering is already done, sitting in a note on your second monitor, three lines long. You say what moved. You say what's next. You say what you need, including the blocker you'd normally swallow. Twenty-five seconds later you're done, and you spend the rest of the call listening instead of recovering.
The engineering is the hard part. This part is just a format nobody bothered to teach you.
If you want to practise this with someone who'll tell you what's actually landing, AptSpeak does 1:1 sessions built around real situations like standup, sprint demos, and pushing back in a design review. Sessions are with a coach who works specifically with developers, so you won't spend time explaining what a PR is. You can book a single session to see whether it's useful.
Key takeaway
Want your team communicating like this with every client? See how AptSpeak works with outsourcing and staffing firms.