AI chat UI in React: streaming messages, tool calls and reasoning
Build an AI chat UI in React: stream tokens with a ReadableStream reader, render reasoning and tool calls, auto scroll, stop and retry. Works with any provider.
By MiniDev23 min read
An AI chat UI in React comes down to three parts: a message list that renders text, reasoning and tool calls as they stream in, a composer that submits on Enter and turns into a stop button while the model is working, and a small hook that reads the response body chunk by chunk with a ReadableStream reader. None of it depends on a specific model provider. This guide builds all three with free MiniDev UI components and plain fetch.
The components and what each one does
| Component | Role | Key props |
|---|---|---|
| ChatThread | Scrollable message log | messages, empty, className |
| MessageBubble | One turn: user capsule, assistant prose, system rule | role, streaming |
| PromptInput | Auto-growing composer, Enter to send | value, onChange, onSubmit, disabled, toolbar |
| ReasoningBlock | Folded model thinking with a live shimmer | active, duration, title, defaultOpen |
| ToolCallCard | One tool invocation with status and output | name, args, status, duration, children |
| StopGenerating | Outline button with a stop icon | onClick |
| SuggestionChips | Starter prompts for the empty state | items, onSelect |
| ModelPicker | Select for the model | models, value, onChange |
npx shadcn@latest add https://ui.minidev.pro/r/chat-thread.json https://ui.minidev.pro/r/prompt-input.json https://ui.minidev.pro/r/reasoning-block.json https://ui.minidev.pro/r/tool-call-card.json https://ui.minidev.pro/r/stop-generating.json https://ui.minidev.pro/r/suggestion-chips.json https://ui.minidev.pro/r/model-picker.json https://ui.minidev.pro/r/retry-block.json
chat-thread brings message-bubble and streaming-cursor with it. The files land in components/ui, so everything below imports from @/components/ui/*. The components style themselves with MiniDev's semantic tokens, so include the token stylesheet in your global CSS once.
Model the message as parts
A modern assistant turn is not one string. It can start with reasoning, call two tools, then write an answer. Store each turn as an ordered list of parts and let rendering decide how each part looks. This also makes the stream easy to apply: every incoming event either appends to the last part or adds a new one.
export type Part =
| { kind: "text"; text: string }
| { kind: "reasoning"; text: string; ms?: number }
| { kind: "tool"; id: string; name: string; status: "running" | "done" | "error"; args?: string; output?: string }
export type ChatMessage = {
id: string
role: "user" | "assistant"
parts: Part[]
status?: "streaming" | "done" | "stopped" | "error"
}
/** What the server sends, one JSON object per line. */
export type ChatEvent =
| { type: "text"; delta: string }
| { type: "reasoning"; delta: string }
| { type: "reasoning-end"; ms: number }
| { type: "tool"; id: string; name: string; status: "running" | "done" | "error"; args?: string; output?: string }
| { type: "error"; message: string }
The wire format is newline-delimited JSON (NDJSON). It is provider-agnostic on purpose: your server route translates whatever your model SDK emits into these five event types, and the client never changes when you switch providers.
Read the stream with a ReadableStream reader
fetch exposes the response body as a ReadableStream of bytes. Read it with getReader(), decode with a TextDecoder in streaming mode, and split on newlines. Network chunks do not respect line boundaries, so keep a buffer and only parse complete lines.
import type { ChatEvent } from "./chat-types"
export async function* readEvents(res: Response): AsyncGenerator<ChatEvent> {
if (!res.body) throw new Error("Response has no body")
const reader = res.body.getReader()
const decoder = new TextDecoder()
let buffer = ""
while (true) {
const { value, done } = await reader.read()
if (done) break
// stream: true keeps multi-byte characters split across chunks intact
buffer += decoder.decode(value, { stream: true })
let nl: number
while ((nl = buffer.indexOf("\n")) !== -1) {
const line = buffer.slice(0, nl).trim()
buffer = buffer.slice(nl + 1)
if (line) yield JSON.parse(line) as ChatEvent
}
}
buffer += decoder.decode()
if (buffer.trim()) yield JSON.parse(buffer) as ChatEvent
}
If your endpoint speaks Server-Sent Events instead, the loop is the same. Split the buffer on blank lines ("\n\n"), strip the data: prefix from each line, and skip comments and [DONE] sentinels.
Applying an event to a message is a pure function. Text and reasoning deltas extend the last part of the same kind; tool events upsert by id so a card moves from running to done in place.
import type { ChatEvent, Part } from "./chat-types"
export function applyEvent(parts: Part[], ev: ChatEvent): Part[] {
const last = parts[parts.length - 1]
switch (ev.type) {
case "text":
if (last?.kind === "text") return [...parts.slice(0, -1), { ...last, text: last.text + ev.delta }]
return [...parts, { kind: "text", text: ev.delta }]
case "reasoning":
if (last?.kind === "reasoning") return [...parts.slice(0, -1), { ...last, text: last.text + ev.delta }]
return [...parts, { kind: "reasoning", text: ev.delta }]
case "reasoning-end":
return parts.map((p) => (p.kind === "reasoning" && p.ms == null ? { ...p, ms: ev.ms } : p))
case "tool": {
const next: Part = { kind: "tool", id: ev.id, name: ev.name, status: ev.status, args: ev.args, output: ev.output }
const i = parts.findIndex((p) => p.kind === "tool" && p.id === ev.id)
return i === -1 ? [...parts, next] : parts.map((p, j) => (j === i ? next : p))
}
default:
return parts
}
}
A useChat hook with stop and retry
The hook owns the message list, one AbortController per request, and a streaming flag. Stop aborts the controller; the partial answer stays on screen marked as stopped. Retry drops everything after the last user message and runs the request again.
"use client"
import * as React from "react"
import type { ChatMessage } from "@/lib/chat-types"
import { readEvents } from "@/lib/read-events"
import { applyEvent } from "@/lib/apply-event"
const toWire = (m: ChatMessage) => ({
role: m.role,
content: m.parts.map((p) => (p.kind === "text" ? p.text : "")).join(""),
})
export function useChat({ endpoint = "/api/chat", model }: { endpoint?: string; model?: string } = {}) {
const [messages, setMessages] = React.useState<ChatMessage[]>([])
const [streaming, setStreaming] = React.useState(false)
const controller = React.useRef<AbortController | null>(null)
const run = React.useCallback(async (history: ChatMessage[]) => {
const id = crypto.randomUUID()
const update = (fn: (m: ChatMessage) => ChatMessage) =>
setMessages((all) => all.map((m) => (m.id === id ? fn(m) : m)))
setMessages([...history, { id, role: "assistant", parts: [], status: "streaming" }])
setStreaming(true)
const ac = new AbortController()
controller.current = ac
try {
const res = await fetch(endpoint, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ model, messages: history.map(toWire) }),
signal: ac.signal,
})
if (!res.ok) throw new Error("HTTP " + res.status)
for await (const ev of readEvents(res)) {
if (ev.type === "error") throw new Error(ev.message)
update((m) => ({ ...m, parts: applyEvent(m.parts, ev) }))
}
update((m) => ({ ...m, status: "done" }))
} catch {
update((m) => ({ ...m, status: ac.signal.aborted ? "stopped" : "error" }))
} finally {
setStreaming(false)
controller.current = null
}
}, [endpoint, model])
const send = (text: string) => {
if (!text.trim() || streaming) return
run([...messages, { id: crypto.randomUUID(), role: "user", parts: [{ kind: "text", text }] }])
}
const stop = () => controller.current?.abort()
const retry = () => {
const i = messages.findLastIndex((m) => m.role === "user")
if (i !== -1 && !streaming) run(messages.slice(0, i + 1))
}
return { messages, streaming, send, stop, retry }
}
React batches the per-chunk state updates, which is fine for typical token rates. If a fast model sends hundreds of tiny chunks per second, collect deltas in a ref and flush them once per requestAnimationFrame so you render at most once per frame.
Render reasoning, tool calls and text
ChatThread takes messages shaped as { id, role, content, streaming }, where content is any React node. Map each part to a component and pass the result as content. Text parts render as inline spans so the StreamingCursor that MessageBubble appends while streaming is true sits right after the last word.
import { ReasoningBlock } from "@/components/ui/reasoning-block"
import { ToolCallCard } from "@/components/ui/tool-call-card"
import { RetryBlock } from "@/components/ui/retry-block"
import type { ChatMessage } from "@/lib/chat-types"
export function renderMessage(m: ChatMessage, onRetry: () => void) {
const live = m.status === "streaming"
return (
<>
{m.parts.map((p, i) => {
if (p.kind === "reasoning")
return (
<ReasoningBlock
key={i}
className="mb-3"
active={live && p.ms == null}
duration={p.ms != null ? Math.max(1, Math.round(p.ms / 1000)) + "s" : undefined}
>
{p.text}
</ReasoningBlock>
)
if (p.kind === "tool")
return (
<ToolCallCard key={i} className="mb-3" name={p.name} args={p.args} status={p.status} defaultOpen={false}>
{p.output}
</ToolCallCard>
)
return <span key={i} className="whitespace-pre-wrap">{p.text}</span>
})}
{m.status === "stopped" ? <p className="mt-2 text-xs text-fg-subtle">Stopped</p> : null}
{m.status === "error" ? <RetryBlock className="mt-3" onRetry={onRetry} /> : null}
</>
)
}
ReasoningBlock is folded by default. While active is true its label shimmers as "Thinking…"; once you pass duration it reads "Thought for 4s". ToolCallCard shows a spinner and "Running" for status="running", a check for done and a cross for error. When children is empty it renders without a disclosure, so a tool that returns nothing does not show an empty panel.
ThinkingBlock is the simpler alternative: a bordered box with a fixed "Thinking" label that mounts its content only when opened. Use it where you do not track timing. StreamingMessage is a static bubble with a cursor, useful for skeleton states and demos.
Reasoning BlockuiCollapsible AI reasoning block, folded by default. The title shimmers while the model is thinking, then reads Thought for Ns, with an animated height reveal.npx shadcn@latest add ui.minidev.pro/r/reasoning-block.jsonTool Call CarduiAgent tool call card showing the tool name, argument preview, duration and a running, done or failed status chip, with collapsible monospace output below.npx shadcn@latest add ui.minidev.pro/r/tool-call-card.jsonPut the page together
The thread must own the scroll, so give it the remaining height in a flex column with min-h-0 flex-1. The composer sits below. While the model is streaming, show StopGenerating above the input and ignore submits; the textarea stays enabled so people can type their next message.
"use client"
import * as React from "react"
import { ChatThread } from "@/components/ui/chat-thread"
import { PromptInput } from "@/components/ui/prompt-input"
import { StopGenerating } from "@/components/ui/stop-generating"
import { SuggestionChips } from "@/components/ui/suggestion-chips"
import { ModelPicker } from "@/components/ui/model-picker"
import { useChat } from "@/hooks/use-chat"
import { renderMessage } from "@/components/chat/render-message"
const MODELS = [{ id: "fast", label: "Fast" }, { id: "deep", label: "Deep reasoning" }]
export default function ChatPage() {
const [model, setModel] = React.useState("fast")
const [draft, setDraft] = React.useState("")
const { messages, streaming, send, stop, retry } = useChat({ model })
return (
<div className="mx-auto flex h-dvh max-w-3xl flex-col">
<ChatThread
className="min-h-0 flex-1 px-4 py-6"
messages={messages.map((m) => ({
id: m.id,
role: m.role,
streaming: m.status === "streaming",
content: renderMessage(m, retry),
}))}
empty={
<div className="grid flex-1 place-items-center content-center gap-4 py-16 text-center">
<p className="text-lg font-medium text-fg">What are we working on?</p>
<SuggestionChips items={["Summarize this PR", "Draft a release note", "Explain this error"]} onSelect={send} />
</div>
}
/>
<div className="space-y-2 p-3">
{streaming ? <div className="flex justify-center"><StopGenerating onClick={stop} /></div> : null}
<PromptInput
value={draft}
onChange={setDraft}
onSubmit={() => {
if (streaming) return
send(draft)
setDraft("")
}}
toolbar={<ModelPicker models={MODELS} value={model} onChange={setModel} className="h-8 w-44" />}
/>
</div>
</div>
)
}
PromptInput grows with its content up to 240px, submits on Enter, inserts a newline on Shift+Enter, and only enables the send button when the trimmed value is non-empty. The toolbar slot sits to the right of the attach button, which is where a model picker or context chips belong.
Auto scroll that respects the reader
Out of the box, ChatThread sets scrollTop = scrollHeight whenever messages changes, which keeps a streaming answer in view. If you want people to be able to scroll up and read while the answer keeps streaming, only follow the bottom when the reader is already near it. Since the file is in your repo, replace its effect:
const stick = React.useRef(true)
const count = React.useRef(0)
React.useEffect(() => {
const el = ref.current
if (!el) return
// A new turn snaps back to the bottom; token updates follow only if the reader is there.
if (messages.length > count.current) stick.current = true
count.current = messages.length
if (stick.current) el.scrollTop = el.scrollHeight
}, [messages])
// on the scrolling div:
onScroll={(e) => {
const el = e.currentTarget
stick.current = el.scrollHeight - el.scrollTop - el.clientHeight < 48
}}
Streaming tokens update the same message, so the list length only grows when a turn starts. That is the signal to snap back down. A "Jump to latest" button that appears while the reader is scrolled away completes the pattern; keep stick in state instead of a ref if you render one.
The thread is a role="log" with aria-live="polite", so screen readers announce new content. Token-level updates can be noisy; setting aria-busy on the log while streaming is true asks assistive technology to wait until the answer settles.
The server route
Keep API keys on the server. A Next.js route handler calls your provider, translates its stream into ChatEvent lines, and forwards req.signal so a stop in the browser also cancels the upstream request.
import { streamModel } from "@/lib/model" // your adapter: yields ChatEvent objects
export async function POST(req: Request) {
const { messages, model } = await req.json()
const encoder = new TextEncoder()
const stream = new ReadableStream({
async start(controller) {
const push = (ev: unknown) => controller.enqueue(encoder.encode(JSON.stringify(ev) + "\n"))
try {
for await (const ev of streamModel({ messages, model, signal: req.signal })) push(ev)
} catch {
if (!req.signal.aborted) push({ type: "error", message: "The model request failed." })
} finally {
try { controller.close() } catch {}
}
},
})
return new Response(stream, {
headers: { "Content-Type": "application/x-ndjson; charset=utf-8", "Cache-Control": "no-store" },
})
}
streamModel is the only provider-specific code in the whole stack: an async generator that calls your SDK and yields { type: "text", delta } for content tokens, reasoning deltas for thinking, and tool events when the model calls and finishes a function.
Why NDJSON over a POST
EventSource only issues GET requests, so it cannot carry a chat history in the body. A plain fetch POST that returns NDJSON can, it works through the same reader loop as SSE, and it needs no client library. Two operational details: make sure nothing between the server and the browser buffers the response (for example, nginx needs X-Accel-Buffering: no), and never cache the route. If the stream arrives in one lump at the end, buffering is almost always the reason.
Finishing touches
- Put MessageActions under finished assistant turns for copy, retry and feedback. Its
onRetrycan call the sameretryfrom the hook. - Render markdown in text parts with a parser of your choice once the stream is complete, and plain text while streaming, to avoid half-open code fences flickering.
- Persist the thread on the server after
statusbecomesdone, not per chunk. - Limit what you send back: trim old turns or summarize them before the history exceeds the model context.
- Errors can arrive mid-answer. The hook keeps the partial text and marks the turn
error, so the reader sees what arrived plus aRetryBlockinstead of losing everything. - Show per-turn usage or cost with TokenUsageMeter when people pay by the token.
Components and install
All of these are free and MIT licensed in MiniDev UI's AI chat set. The SaaS dashboard guide shows how to host a chat panel inside an app shell, and the MiniDev studio builds complete AI products on the same kit if you want a team to ship it.
Chat ThreaduiScrolling chat thread for React that renders user, assistant and system messages as bubbles, auto scrolls to the newest one and uses a live log role.npx shadcn@latest add ui.minidev.pro/r/chat-thread.jsonPrompt InputuiAn AI chat composer that grows with its content, sends on Enter, adds a newline on Shift Enter and has attach, toolbar and send controls. Built for React.npx shadcn@latest add ui.minidev.pro/r/prompt-input.jsonStop GeneratinguiStop generating button with a filled square icon and an onClick handler, for cancelling a streaming AI response mid-answer in chat interfaces built with React.npx shadcn@latest add ui.minidev.pro/r/stop-generating.jsonSuggestion ChipsuiRow of rounded suggestion chips such as Summarize or Find bugs that call onSelect when clicked. Use them as prompt starters under a React AI chat input.npx shadcn@latest add ui.minidev.pro/r/suggestion-chips.jsonnpx shadcn@latest add https://ui.minidev.pro/r/<name>.json
Components used in this guide
Frequently asked questions
How do I stream a chat response in React without a provider SDK?
Call fetch, read res.body.getReader() in a loop, decode each chunk with TextDecoder using { stream: true }, and split complete lines out of a buffer. Update the assistant message in state as each event arrives.
How do I add a stop button to a streaming chat?
Create an AbortController per request, pass its signal to fetch, and call abort() from the stop button. The read loop throws, and you mark the partial message as stopped instead of failed by checking signal.aborted.
How should tool calls appear in a chat UI?
As compact cards inside the assistant turn that show the tool name, a short argument preview and a running, done or failed status, with the output collapsed. Upsert them by call id so the same card updates in place.
Can I build a ChatGPT style UI with shadcn components?
Yes. MiniDev UI is a shadcn-compatible registry, so npx shadcn@latest add https://ui.minidev.pro/r/chat-thread.json installs the thread, bubble and cursor into your project alongside any shadcn components you already use.