3.1 THE APPLICATION

Chat UX

Input modes, thinking indicators, tool exposure, and how the interface changes when the model is taking actions, not just answering questions.

AUTHOR LEA DANG WORDS 3,473 READ 15 MIN SECTIONS 15

A chat box is the wrong default. Most LLM products started as chatbots and then bolted agents on underneath, which means the same UI is now doing two jobs that have almost nothing in common. A chatbot is a turn-based conversation. An agent is a loop: tool calls, partial output, multi-step actions, sometimes minutes of work between user messages.

Every chat UX problem gets harder when the thing on the other side is an agent. Latency stops being measured in seconds and starts being measured in minutes. “Thinking” stops being a poetic euphemism for inference and becomes a real fact about the system. Tool calls go from invisible plumbing to the main content the user is paying attention to. Stop, edit, and regenerate stop being convenience features and become the only way to steer a system that’s actively doing things.

This chapter is about that layer. The rules from plain chat mostly still apply, they just need updating once the model is taking actions instead of only answering.

The five patterns

Five patterns are doing essentially all the work in production right now. Most apps mix two or three, and knowing which one fits which moment is most of the design problem.

Pick the wrong one and the product feels off even when the model is fine. A chat app with skeletons looks like a form. A voice agent with no acknowledgment phrases feels rude. The rest of this chapter is about executing each of these well, and switching between them when you should.

Input modalities

Plain text typing is the default and almost always the right default. A handful of alternatives have shown up in shipped agent products and meaningfully change what the conversation feels like.

Text typing

The boring baseline. Free-form, editable, scrolls back. Users can revise before sending. The downside is that typing is slow and kills momentum, especially on mobile, especially when the user is thinking out loud.

Voice-to-text, push-to-talk (Claude voice input, Wispr Flow)

User holds a key, speaks, the system transcribes, the transcription is inserted into the composer as editable text. User reads it back, fixes the obvious mistakes, hits send. This is the mode Claude exposes for voice input, and the mode Wispr Flow built an entire product around.

  • Pros: much faster than typing for long inputs. Captures intent fully formed because the user is talking instead of constructing a sentence under cognitive load. Stays compatible with everything else in the chat UI. Still text, still editable, still searchable.
  • Cons: still synchronous, you wait for transcription before sending. Adds a “review” step that becomes friction when the transcript is wrong. Doesn’t reduce model latency at all, just input latency.

Real-time voice (ChatGPT Advanced Voice, OpenAI Realtime, Gemini Live)

Speech to speech. The user talks, the model talks back, both in real time, with interruption. Built on dedicated real-time APIs that bypass the usual transcribe-then-generate pipeline.

  • Pros: the lowest TTFT possible. The conversation feels like a conversation, not a transaction. Interruption works the way it works between humans, which fixes “stuck waiting for the wrong answer” with no UI at all.
  • Cons: by far the hardest mode to build, which is why it’s the rarest in production despite being the most magical when it works. No scrollback by default, so users lose context. Useless in public spaces. The model has to manage turn-taking under uncertainty. Editing what you said is impossible.

Multimodal attach (Claude images, Cursor code context, ChatGPT files)

Drag a file, screenshot, or image into the composer. The agent reads it as part of the prompt.

  • Pros: avoids verbalizing the obvious. “Fix this bug” plus a screenshot of the error is faster than describing the error in words.
  • Cons: users forget the feature exists. Discoverability needs a visible affordance in the composer, not just drag and drop.

Slash commands and forms (Claude Code, Replit Agent)

Structured inputs. /review opens a review form. /test runs the test suite. Slots, dropdowns, parameters.

  • Pros: deterministic, low ambiguity. Removes the “did the model understand?” doubt entirely.
  • Cons: kills the conversational feel. Better thought of as a power-user shortcut layered on top of chat, not a replacement.

The first 200ms

Show something within 200ms

Within 200ms of the user hitting send, something on screen has to change. Otherwise their brain pings the “did the click work” reflex and you’ve lost trust before the agent has done anything.

Cheap things that count:

  • Move the user’s message into the transcript instantly.
  • Drop three dots in the agent slot.
  • Print a status line. “Thinking…”, “Searching the web…”, “Reading 4 files…”.

200ms is the human threshold for “did the click do anything.” Below it, the interaction feels caused by the click. Above it, the user starts wondering, and wondering is where you lose them.

Render before the network

The user pressing send is a UI event, not a network event. Optimistically render their message and an empty agent slot before any byte leaves the browser. If the request fails, swap the slot for an error state. You still hit the 200ms promise, and you don’t owe the API anything for it.

The states of a turn

Most janky chat UIs treat “loading” as a single flag. Users can absolutely tell the difference between an agent that is thinking, typing, calling a tool, stuck on the network, or quietly on fire. Collapse those into one spinner and they have no idea whether to wait, retry, or close the tab.

FIG_001[ ONE TURN ]
IDLESUBMITTINGREASONINGUSING TOOLSTREAMINGCOMPLETEloopsERROR(from any state)failINTERRUPTEDstop
The state machine of a single agent turn. Reasoning and using-tool are where the time goes.

The minimum viable state machine for a single agent turn has eight states, not one:

  1. idle. Composer ready, nothing in flight.
  2. submitting. Request in flight, no tokens yet.
  3. reasoning. Model is thinking, no user-visible text yet.
  4. using_tool. A tool call is running. Show the tool’s name.
  5. streaming. Tokens are landing in the bubble.
  6. complete. Done. Re-enable the composer.
  7. interrupted. User hit stop. Keep the partial response.
  8. error. Show what failed. Offer a retry.

One face per state

Submitting is dots. Reasoning is a labeled collapsible. Using_tool is a status line with the tool name. Streaming is the live cursor. The user reads the state off the visual without thinking about it, which is the entire point.

Thinking indicators

Dots

Dots (or a pulsing avatar) say “the system is alive, no idea how long this will take.” Cheap, honest, fine when you genuinely don’t know the latency. After about three seconds with nothing else on screen, they start to feel like a hang. Don’t ship dots alone past five.

Status text

“Searching for…”, “Reading 4 files…”, “Drafting…”. The right move for any tool-using agent. The text is not really a loading indicator. It IS the engagement. Users will happily watch “Reading files” for twenty seconds because they understand what’s happening.

One rule: don’t lie. If the label says “Searching” and the backend is actually waiting in a queue, the user notices the dissonance even if they can’t articulate it.

Reasoning blocks

The o1 / o3 / extended-thinking pattern. The model emits chain-of-thought tokens into a collapsible region. Users can pop it open to follow the reasoning or ignore it.

The trap is leakage. Raw reasoning often contains false starts, weird speculation, and the occasional brand-off-message phrasing the model would never actually output as an answer. Show summaries by default. Surface the raw stream only behind a click, and only if you’re confident your model is safe to read out loud.

Where status text comes from

The “Searching the web…” line has to come from somewhere. There are three sensible sources, and you should be intentional about which one you’re shipping.

Hardcoded templates

Each tool call is mapped to a pre-canned phrase. web_search becomes “Searching the web.” read_file becomes “Reading the file.” Arguments can be interpolated in.

  • Pros: instant. No extra inference, deterministic, zero leakage risk, no chance of the model claiming to be doing something it isn’t.
  • Cons: generic. After the third “Searching the web” in a row, the user stops reading. The status text loses signal.

LLM-generated preambles

Before each tool call, the model emits a short rationale. “Looking up the GitHub Actions matrix syntax.” Sometimes this is a natural by-product of chain-of-thought, sometimes you prompt for it explicitly.

  • Pros: specific and contextual. Feels intelligent because it IS.
  • Cons: adds a generation step, which costs tokens and adds latency. The model can claim to be doing things it isn’t. False starts can leak into the UI if you forget to filter.

Hybrid

Tool names and arguments come from the hardcoded template (instant, accurate, faithful to what’s actually happening). One generated sentence per logical chunk of work: a plan at the start, a short narration between phases, a one-line summary at the end. Deterministic stuff stays deterministic, contextual stuff stays contextual.

This is roughly what Claude Code does. Tool call rows render straight from the structured tool-call event the moment it arrives. Between blocks of tool calls, the model emits a short narration (“I’ll start by reading the spec, then check the test file.”). The user gets both the work and the reasoning, each from the source it can be trusted from.

Exposing tool work

This is the question agent products spend the most time on and get wrong most often. How much of what the agent is doing should the user actually see?

Full transparency (coding agents)

Claude Code, Cursor, Aider, and Devin render every tool call. Name, arguments, status, result. Diffs inline. It looks more like a terminal trace than a chat bubble. For users who can read code, this IS the engagement: watching the agent grep, read, edit, test. You can audit every step. You can interrupt at the right moment because you can see what’s about to happen.

  • Pros: maximum trust, full auditability, debugging is possible at all.
  • Cons: visually overwhelming for non-technical users. Slow to scroll through. The actual answer can get lost in the trace.

Aggregated cards (consumer apps)

Perplexity, ChatGPT browsing, and Claude.ai’s web search all collapse multiple tool calls into a single summary card. “Searched 4 sources.” Source chips at the bottom. The work itself is hidden behind a “Show steps” disclosure. Output is the focus, work is implied.

  • Pros: clean, scannable, friendly for non-technical users.
  • Cons: harder to trust without clicking through. You can’t tell which source backed which claim. Hard to debug when the result looks wrong.

The middle ground

One visible row per logical step: “Reading 4 files.” Click expands the details. Top-level facts inline (the search query, the file count). Deep work (raw HTML, full file contents) hidden behind disclosure. Default closed for consumer apps, default open for developer tools.

Two specific rules worth following:

  • Surface failure, hide success. A successful read_file doesn’t need its 800-line output displayed. A failing tool call should be loud, with the error visible without clicking.
  • Group consecutive calls. Six read_file calls in a row should render as one “Reading 6 files” row, not six separate rows. The user’s attention budget is limited and you should spend it on novelty.

How ChatGPT shows reasoning

GPT-5.4 Thinking, o3, and o4-mini have converged on a four-part shape.

A short preamble lands inside a second. “I’ll check the spec, then compare it to the implementation.” Buys patience for the slower thinking that follows.

A collapsible thinking block with a live elapsed-time counter. The label rotates (“Analyzing constraints”, “Comparing options”) as the model moves through steps. The counter is a quiet honesty contract: the user can see how long they’ve been waiting and decide if it’s worth it.

Mid-thought interjection. While reasoning is still streaming, the composer stays open. The user can append guidance and the model folds it in before answering. This is the cleanest fix in the wild for the worst case of long thinking, which is the user noticing the model is going the wrong direction and being stuck.

The final answer streams as normal text. The thinking block auto-collapses on completion but stays expandable.

The animated icon, the rotating verb, and the counter each do little on their own. Together they create the impression of progress where no progress bar exists.

How Claude shows tool use

Same bones, different emphasis. Claude leans harder into making the work visible.

Thinking ships as a separate event type in the API response, not interleaved with user-visible text. The frontend decides what to show. If your frontend assumes every streaming delta is user text, you will dump the model’s internal speculation straight into the UI. Filter on event type. It is the most important defensive line you’ll write.

Tool use renders inline. Every call shows up as a structured row: name, brief argument summary, status (running / done / failed), one-line result. It works because the engagement IS the visibility of work happening.

Artifacts in claude.ai render long structured outputs (code, documents, HTML) in a side panel instead of scrolling inside the chat. The chat shows a small “Created X” card. The streaming-markdown problem mostly goes away for big outputs, and the user gets a stable artifact to come back to.

The two-tier preface

If you ship on a model with slow TTFT, this is the highest-leverage move available to you. A small fast model emits a one-line preface almost immediately while the main model does the real work, and you stitch the two into one stream.

FIG_002[ TWO-TIER PREFACE ]
USERFAST MODELpreface · 200-800msMAIN MODELanswer · secs to minsONE STREAM
A fast model covers the wait while the main model does the real work.

The user sees activity at human speed. The main agent gets its four to eight seconds (or four to eight minutes) without anyone noticing. This is roughly what ChatGPT Advanced Voice and Gemini Live do under the hood: conversational filler from a small model, substantive answer from a big one, stitched together as if it were always one stream.

It bites you in two ways:

The preface contradicts the answer. Fast model says “Sure, that’s easy.” Main model says “Actually this is harder than it looks.” Keep the preface tonal, not committal. “Let me check” is safer than “Sure thing.”

The preface sounds like a different person. Different model, different voice. Tune the small model with a system prompt that mirrors the big one. Otherwise users feel the seam.

Typing during a response

What happens when the user fires off another message before the agent is done? Three legit answers, each with a tradeoff.

Block

Disable the composer until the response completes. Simplest, hardest to break, most frustrating for power users. Default for v0 of any chat product.

Cancel and restart

Treat the second message as a hard interrupt. Stop the current generation, drop the partial (or fold it into the new turn’s context), start over. This is what most voice products do, because interrupting someone is socially fine when they’re talking but feels rude in text.

Queue

Accept the new message immediately. Show it in the transcript with a “queued” indicator. Submit it as soon as the current turn finishes. The next turn’s context then includes the original user message, the agent’s full response, and the queued one.

Queuing is the kindest and the fiddliest. The queued message looks present, so users assume it’s already being processed, so they’re confused when the agent only answers the first. Distinguish queued bubbles visually: muted, with a “queued” label that brightens up when its turn actually comes.

Two ways to ruin it

Silently dropping the second message. Users lose work and stop trusting the composer.

Concatenating the queued message into the in-flight request. At best you confuse the agent. At worst you scramble your tool-call audit trail and your traces look like the model was hallucinating user input.

Stop, regenerate, edit

These three buttons do most of the work of unblocking a conversation that’s gone off.

Stop

The stop button has to propagate to the server and abort the agent loop, not just halt rendering on the client. A “stop” that doesn’t stop burns tokens and tool calls the user can’t see, and produces a quiet billing surprise later. Every major SDK supports request abortion. Wire it through.

Regenerate

Re-runs the last turn with the same input. The design question is whether to discard the previous attempt or keep it as a sibling branch. ChatGPT keeps regenerations accessible via arrows. Claude.ai keeps history linear. Branching is more powerful and more confusing. Pick one consciously and stop second-guessing.

Edit

Lets the user modify an earlier message and re-run from there. The most powerful repair tool in any LLM product. When a conversation has drifted, editing the message that started the drift beats asking the agent politely to please back up.

The gotcha: edits must truncate the conversation at that point. Keep the stale message and its response in context, and your follow-ups will be incoherent because the model can see both the old and the new version of what the user wanted.

Streaming gotchas

Things that look easy and are not.

Partial markdown

Partial markdown is ugly. A half-written code block renders as an unclosed styled fragment. A half-written table is a misaligned crime scene. Parse the stream into a markdown AST incrementally. Defer rendering of any unclosed block-level element. Render its raw text in a neutral container, then promote it when the closing fence or row arrives.

Auto-scroll

Naive auto-scroll pins the viewport to the bottom on every chunk. The moment a user scrolls up to re-read something, they get yanked back down. Rule: auto-scroll only if the user was already at the bottom before the chunk arrived. Otherwise show a “jump to latest” button and leave them where they are.

Tool calls in invocation order

If you render tool calls as separate cards, render them in invocation order, not completion order. Async tools can return out of sequence, and if you key by tool-call ID and render in arrival order, the trace looks reordered to the user. Sort by when the model decided to call it, not when the call returned.

Kill the cursor

A blinking cursor at the end of a finished message reads as “still typing.” Once the stream completes, kill it. Half a second of extra blink makes a finished turn feel unfinished, and the user waits for tokens that aren’t coming.

When the wait is minutes

Some agent workflows take minutes. A coding agent grinding through a long refactor. A research agent reading thirty sources. Streaming UX falls apart at that scale, because nobody watches a status line for four minutes.

Switch modes past 30 seconds

The chat stops being a conversation surface and becomes a job-launch surface.

  • Agent replies with a short confirmation (“Started, I’ll let you know in a few minutes”) and a job card.
  • The job card shows progress in place: tool calls completed, current step, elapsed time.
  • On completion, the result lands as a normal agent message and a system notification. Browser, email, Slack, whatever fits the product.
  • The user can navigate away and come back. Closing the tab does not cancel the job.

What the job card shows

At minimum: current step, elapsed time, a stop button, and one specific recent fact (“Just finished reading auth.ts”). The recent fact is what keeps the card from feeling like a frozen progress bar with extra steps.

Claude Code’s background tasks, ChatGPT’s deep research mode, and most agent platforms have converged on roughly this shape. Chat UI handles sub-30-second interactions. Anything longer wants a different container.

The one metric to instrument

If you instrument exactly one number for agent chat, instrument TTFT, p50 and p95, segmented by route.

Routes below the perception threshold are working. Routes that creep up are where users start abandoning, often weeks before you would have noticed from any other dashboard. Every trick in this chapter is in service of pulling that number down, or, when you genuinely can’t, of giving the user something honest and specific to look at while they wait.