AI chat UI in React: streaming messages, tool calls and reasoning

Build an AI chat UI in React: stream tokens with a ReadableStream reader, render reasoning and tool calls, auto scroll, stop and retry. Works with any provider.

By MiniDev23 min read

An AI chat UI in React comes down to three parts: a message list that renders text, reasoning and tool calls as they stream in, a composer that submits on Enter and turns into a stop button while the model is working, and a small hook that reads the response body chunk by chunk with a ReadableStream reader. None of it depends on a specific model provider. This guide builds all three with free MiniDev UI components and plain fetch.

The components and what each one does

ComponentRoleKey props
ChatThreadScrollable message logmessages, empty, className
MessageBubbleOne turn: user capsule, assistant prose, system rulerole, streaming
PromptInputAuto-growing composer, Enter to sendvalue, onChange, onSubmit, disabled, toolbar
ReasoningBlockFolded model thinking with a live shimmeractive, duration, title, defaultOpen
ToolCallCardOne tool invocation with status and outputname, args, status, duration, children
StopGeneratingOutline button with a stop icononClick
SuggestionChipsStarter prompts for the empty stateitems, onSelect
ModelPickerSelect for the modelmodels, value, onChange
bash
npx shadcn@latest add https://ui.minidev.pro/r/chat-thread.json https://ui.minidev.pro/r/prompt-input.json https://ui.minidev.pro/r/reasoning-block.json https://ui.minidev.pro/r/tool-call-card.json https://ui.minidev.pro/r/stop-generating.json https://ui.minidev.pro/r/suggestion-chips.json https://ui.minidev.pro/r/model-picker.json https://ui.minidev.pro/r/retry-block.json

chat-thread brings message-bubble and streaming-cursor with it. The files land in components/ui, so everything below imports from @/components/ui/*. The components style themselves with MiniDev's semantic tokens, so include the token stylesheet in your global CSS once.

Model the message as parts

A modern assistant turn is not one string. It can start with reasoning, call two tools, then write an answer. Store each turn as an ordered list of parts and let rendering decide how each part looks. This also makes the stream easy to apply: every incoming event either appends to the last part or adds a new one.

lib/chat-types.tsts
export type Part =
  | { kind: "text"; text: string }
  | { kind: "reasoning"; text: string; ms?: number }
  | { kind: "tool"; id: string; name: string; status: "running" | "done" | "error"; args?: string; output?: string }

export type ChatMessage = {
  id: string
  role: "user" | "assistant"
  parts: Part[]
  status?: "streaming" | "done" | "stopped" | "error"
}

/** What the server sends, one JSON object per line. */
export type ChatEvent =
  | { type: "text"; delta: string }
  | { type: "reasoning"; delta: string }
  | { type: "reasoning-end"; ms: number }
  | { type: "tool"; id: string; name: string; status: "running" | "done" | "error"; args?: string; output?: string }
  | { type: "error"; message: string }

The wire format is newline-delimited JSON (NDJSON). It is provider-agnostic on purpose: your server route translates whatever your model SDK emits into these five event types, and the client never changes when you switch providers.

Read the stream with a ReadableStream reader

fetch exposes the response body as a ReadableStream of bytes. Read it with getReader(), decode with a TextDecoder in streaming mode, and split on newlines. Network chunks do not respect line boundaries, so keep a buffer and only parse complete lines.

lib/read-events.tsts
import type { ChatEvent } from "./chat-types"

export async function* readEvents(res: Response): AsyncGenerator<ChatEvent> {
  if (!res.body) throw new Error("Response has no body")
  const reader = res.body.getReader()
  const decoder = new TextDecoder()
  let buffer = ""
  while (true) {
    const { value, done } = await reader.read()
    if (done) break
    // stream: true keeps multi-byte characters split across chunks intact
    buffer += decoder.decode(value, { stream: true })
    let nl: number
    while ((nl = buffer.indexOf("\n")) !== -1) {
      const line = buffer.slice(0, nl).trim()
      buffer = buffer.slice(nl + 1)
      if (line) yield JSON.parse(line) as ChatEvent
    }
  }
  buffer += decoder.decode()
  if (buffer.trim()) yield JSON.parse(buffer) as ChatEvent
}

If your endpoint speaks Server-Sent Events instead, the loop is the same. Split the buffer on blank lines ("\n\n"), strip the data: prefix from each line, and skip comments and [DONE] sentinels.

Applying an event to a message is a pure function. Text and reasoning deltas extend the last part of the same kind; tool events upsert by id so a card moves from running to done in place.

lib/apply-event.tsts
import type { ChatEvent, Part } from "./chat-types"

export function applyEvent(parts: Part[], ev: ChatEvent): Part[] {
  const last = parts[parts.length - 1]
  switch (ev.type) {
    case "text":
      if (last?.kind === "text") return [...parts.slice(0, -1), { ...last, text: last.text + ev.delta }]
      return [...parts, { kind: "text", text: ev.delta }]
    case "reasoning":
      if (last?.kind === "reasoning") return [...parts.slice(0, -1), { ...last, text: last.text + ev.delta }]
      return [...parts, { kind: "reasoning", text: ev.delta }]
    case "reasoning-end":
      return parts.map((p) => (p.kind === "reasoning" && p.ms == null ? { ...p, ms: ev.ms } : p))
    case "tool": {
      const next: Part = { kind: "tool", id: ev.id, name: ev.name, status: ev.status, args: ev.args, output: ev.output }
      const i = parts.findIndex((p) => p.kind === "tool" && p.id === ev.id)
      return i === -1 ? [...parts, next] : parts.map((p, j) => (j === i ? next : p))
    }
    default:
      return parts
  }
}

A useChat hook with stop and retry

The hook owns the message list, one AbortController per request, and a streaming flag. Stop aborts the controller; the partial answer stays on screen marked as stopped. Retry drops everything after the last user message and runs the request again.

hooks/use-chat.tstsx
"use client"
import * as React from "react"
import type { ChatMessage } from "@/lib/chat-types"
import { readEvents } from "@/lib/read-events"
import { applyEvent } from "@/lib/apply-event"

const toWire = (m: ChatMessage) => ({
  role: m.role,
  content: m.parts.map((p) => (p.kind === "text" ? p.text : "")).join(""),
})

export function useChat({ endpoint = "/api/chat", model }: { endpoint?: string; model?: string } = {}) {
  const [messages, setMessages] = React.useState<ChatMessage[]>([])
  const [streaming, setStreaming] = React.useState(false)
  const controller = React.useRef<AbortController | null>(null)

  const run = React.useCallback(async (history: ChatMessage[]) => {
    const id = crypto.randomUUID()
    const update = (fn: (m: ChatMessage) => ChatMessage) =>
      setMessages((all) => all.map((m) => (m.id === id ? fn(m) : m)))
    setMessages([...history, { id, role: "assistant", parts: [], status: "streaming" }])
    setStreaming(true)
    const ac = new AbortController()
    controller.current = ac
    try {
      const res = await fetch(endpoint, {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ model, messages: history.map(toWire) }),
        signal: ac.signal,
      })
      if (!res.ok) throw new Error("HTTP " + res.status)
      for await (const ev of readEvents(res)) {
        if (ev.type === "error") throw new Error(ev.message)
        update((m) => ({ ...m, parts: applyEvent(m.parts, ev) }))
      }
      update((m) => ({ ...m, status: "done" }))
    } catch {
      update((m) => ({ ...m, status: ac.signal.aborted ? "stopped" : "error" }))
    } finally {
      setStreaming(false)
      controller.current = null
    }
  }, [endpoint, model])

  const send = (text: string) => {
    if (!text.trim() || streaming) return
    run([...messages, { id: crypto.randomUUID(), role: "user", parts: [{ kind: "text", text }] }])
  }
  const stop = () => controller.current?.abort()
  const retry = () => {
    const i = messages.findLastIndex((m) => m.role === "user")
    if (i !== -1 && !streaming) run(messages.slice(0, i + 1))
  }

  return { messages, streaming, send, stop, retry }
}

React batches the per-chunk state updates, which is fine for typical token rates. If a fast model sends hundreds of tiny chunks per second, collect deltas in a ref and flush them once per requestAnimationFrame so you render at most once per frame.

Render reasoning, tool calls and text

ChatThread takes messages shaped as { id, role, content, streaming }, where content is any React node. Map each part to a component and pass the result as content. Text parts render as inline spans so the StreamingCursor that MessageBubble appends while streaming is true sits right after the last word.

components/chat/render-message.tsxtsx
import { ReasoningBlock } from "@/components/ui/reasoning-block"
import { ToolCallCard } from "@/components/ui/tool-call-card"
import { RetryBlock } from "@/components/ui/retry-block"
import type { ChatMessage } from "@/lib/chat-types"

export function renderMessage(m: ChatMessage, onRetry: () => void) {
  const live = m.status === "streaming"
  return (
    <>
      {m.parts.map((p, i) => {
        if (p.kind === "reasoning")
          return (
            <ReasoningBlock
              key={i}
              className="mb-3"
              active={live && p.ms == null}
              duration={p.ms != null ? Math.max(1, Math.round(p.ms / 1000)) + "s" : undefined}
            >
              {p.text}
            </ReasoningBlock>
          )
        if (p.kind === "tool")
          return (
            <ToolCallCard key={i} className="mb-3" name={p.name} args={p.args} status={p.status} defaultOpen={false}>
              {p.output}
            </ToolCallCard>
          )
        return <span key={i} className="whitespace-pre-wrap">{p.text}</span>
      })}
      {m.status === "stopped" ? <p className="mt-2 text-xs text-fg-subtle">Stopped</p> : null}
      {m.status === "error" ? <RetryBlock className="mt-3" onRetry={onRetry} /> : null}
    </>
  )
}

ReasoningBlock is folded by default. While active is true its label shimmers as "Thinking…"; once you pass duration it reads "Thought for 4s". ToolCallCard shows a spinner and "Running" for status="running", a check for done and a cross for error. When children is empty it renders without a disclosure, so a tool that returns nothing does not show an empty panel.

ThinkingBlock is the simpler alternative: a bordered box with a fixed "Thinking" label that mounts its content only when opened. Use it where you do not track timing. StreamingMessage is a static bubble with a cursor, useful for skeleton states and demos.

Reasoning BlockuiCollapsible AI reasoning block, folded by default. The title shimmers while the model is thinking, then reads Thought for Ns, with an animated height reveal.npx shadcn@latest add ui.minidev.pro/r/reasoning-block.jsonTool Call CarduiAgent tool call card showing the tool name, argument preview, duration and a running, done or failed status chip, with collapsible monospace output below.npx shadcn@latest add ui.minidev.pro/r/tool-call-card.json

Put the page together

The thread must own the scroll, so give it the remaining height in a flex column with min-h-0 flex-1. The composer sits below. While the model is streaming, show StopGenerating above the input and ignore submits; the textarea stays enabled so people can type their next message.

app/chat/page.tsxtsx
"use client"
import * as React from "react"
import { ChatThread } from "@/components/ui/chat-thread"
import { PromptInput } from "@/components/ui/prompt-input"
import { StopGenerating } from "@/components/ui/stop-generating"
import { SuggestionChips } from "@/components/ui/suggestion-chips"
import { ModelPicker } from "@/components/ui/model-picker"
import { useChat } from "@/hooks/use-chat"
import { renderMessage } from "@/components/chat/render-message"

const MODELS = [{ id: "fast", label: "Fast" }, { id: "deep", label: "Deep reasoning" }]

export default function ChatPage() {
  const [model, setModel] = React.useState("fast")
  const [draft, setDraft] = React.useState("")
  const { messages, streaming, send, stop, retry } = useChat({ model })

  return (
    <div className="mx-auto flex h-dvh max-w-3xl flex-col">
      <ChatThread
        className="min-h-0 flex-1 px-4 py-6"
        messages={messages.map((m) => ({
          id: m.id,
          role: m.role,
          streaming: m.status === "streaming",
          content: renderMessage(m, retry),
        }))}
        empty={
          <div className="grid flex-1 place-items-center content-center gap-4 py-16 text-center">
            <p className="text-lg font-medium text-fg">What are we working on?</p>
            <SuggestionChips items={["Summarize this PR", "Draft a release note", "Explain this error"]} onSelect={send} />
          </div>
        }
      />
      <div className="space-y-2 p-3">
        {streaming ? <div className="flex justify-center"><StopGenerating onClick={stop} /></div> : null}
        <PromptInput
          value={draft}
          onChange={setDraft}
          onSubmit={() => {
            if (streaming) return
            send(draft)
            setDraft("")
          }}
          toolbar={<ModelPicker models={MODELS} value={model} onChange={setModel} className="h-8 w-44" />}
        />
      </div>
    </div>
  )
}

PromptInput grows with its content up to 240px, submits on Enter, inserts a newline on Shift+Enter, and only enables the send button when the trimmed value is non-empty. The toolbar slot sits to the right of the attach button, which is where a model picker or context chips belong.

Auto scroll that respects the reader

Out of the box, ChatThread sets scrollTop = scrollHeight whenever messages changes, which keeps a streaming answer in view. If you want people to be able to scroll up and read while the answer keeps streaming, only follow the bottom when the reader is already near it. Since the file is in your repo, replace its effect:

components/ui/chat-thread.tsxtsx
const stick = React.useRef(true)
const count = React.useRef(0)

React.useEffect(() => {
  const el = ref.current
  if (!el) return
  // A new turn snaps back to the bottom; token updates follow only if the reader is there.
  if (messages.length > count.current) stick.current = true
  count.current = messages.length
  if (stick.current) el.scrollTop = el.scrollHeight
}, [messages])

// on the scrolling div:
onScroll={(e) => {
  const el = e.currentTarget
  stick.current = el.scrollHeight - el.scrollTop - el.clientHeight < 48
}}

Streaming tokens update the same message, so the list length only grows when a turn starts. That is the signal to snap back down. A "Jump to latest" button that appears while the reader is scrolled away completes the pattern; keep stick in state instead of a ref if you render one.

The thread is a role="log" with aria-live="polite", so screen readers announce new content. Token-level updates can be noisy; setting aria-busy on the log while streaming is true asks assistive technology to wait until the answer settles.

The server route

Keep API keys on the server. A Next.js route handler calls your provider, translates its stream into ChatEvent lines, and forwards req.signal so a stop in the browser also cancels the upstream request.

app/api/chat/route.tsts
import { streamModel } from "@/lib/model" // your adapter: yields ChatEvent objects

export async function POST(req: Request) {
  const { messages, model } = await req.json()
  const encoder = new TextEncoder()
  const stream = new ReadableStream({
    async start(controller) {
      const push = (ev: unknown) => controller.enqueue(encoder.encode(JSON.stringify(ev) + "\n"))
      try {
        for await (const ev of streamModel({ messages, model, signal: req.signal })) push(ev)
      } catch {
        if (!req.signal.aborted) push({ type: "error", message: "The model request failed." })
      } finally {
        try { controller.close() } catch {}
      }
    },
  })
  return new Response(stream, {
    headers: { "Content-Type": "application/x-ndjson; charset=utf-8", "Cache-Control": "no-store" },
  })
}

streamModel is the only provider-specific code in the whole stack: an async generator that calls your SDK and yields { type: "text", delta } for content tokens, reasoning deltas for thinking, and tool events when the model calls and finishes a function.

Why NDJSON over a POST

EventSource only issues GET requests, so it cannot carry a chat history in the body. A plain fetch POST that returns NDJSON can, it works through the same reader loop as SSE, and it needs no client library. Two operational details: make sure nothing between the server and the browser buffers the response (for example, nginx needs X-Accel-Buffering: no), and never cache the route. If the stream arrives in one lump at the end, buffering is almost always the reason.

Finishing touches

  • Put MessageActions under finished assistant turns for copy, retry and feedback. Its onRetry can call the same retry from the hook.
  • Render markdown in text parts with a parser of your choice once the stream is complete, and plain text while streaming, to avoid half-open code fences flickering.
  • Persist the thread on the server after status becomes done, not per chunk.
  • Limit what you send back: trim old turns or summarize them before the history exceeds the model context.
  • Errors can arrive mid-answer. The hook keeps the partial text and marks the turn error, so the reader sees what arrived plus a RetryBlock instead of losing everything.
  • Show per-turn usage or cost with TokenUsageMeter when people pay by the token.

Components and install

All of these are free and MIT licensed in MiniDev UI's AI chat set. The SaaS dashboard guide shows how to host a chat panel inside an app shell, and the MiniDev studio builds complete AI products on the same kit if you want a team to ship it.

Chat ThreaduiScrolling chat thread for React that renders user, assistant and system messages as bubbles, auto scrolls to the newest one and uses a live log role.npx shadcn@latest add ui.minidev.pro/r/chat-thread.jsonPrompt InputuiAn AI chat composer that grows with its content, sends on Enter, adds a newline on Shift Enter and has attach, toolbar and send controls. Built for React.npx shadcn@latest add ui.minidev.pro/r/prompt-input.jsonStop GeneratinguiStop generating button with a filled square icon and an onClick handler, for cancelling a streaming AI response mid-answer in chat interfaces built with React.npx shadcn@latest add ui.minidev.pro/r/stop-generating.jsonSuggestion ChipsuiRow of rounded suggestion chips such as Summarize or Find bugs that call onSelect when clicked. Use them as prompt starters under a React AI chat input.npx shadcn@latest add ui.minidev.pro/r/suggestion-chips.json
bash
npx shadcn@latest add https://ui.minidev.pro/r/<name>.json

Components used in this guide

Chat ThreaduiScrolling chat thread for React that renders user, assistant and system messages as bubbles, auto scrolls to the newest one and uses a live log role.npx shadcn@latest add ui.minidev.pro/r/chat-thread.jsonMessage BubbleuiA React chat message for user, assistant and system roles: user turns sit in a right aligned capsule, assistant replies read as prose, with a streaming cursor.npx shadcn@latest add ui.minidev.pro/r/message-bubble.jsonStreaming MessageuiChat message bubble that renders partial AI output with a blinking caret at the end, for showing token-by-token streaming responses in a Tailwind chat UI.npx shadcn@latest add ui.minidev.pro/r/streaming-message.jsonPrompt InputuiAn AI chat composer that grows with its content, sends on Enter, adds a newline on Shift Enter and has attach, toolbar and send controls. Built for React.npx shadcn@latest add ui.minidev.pro/r/prompt-input.jsonReasoning BlockuiCollapsible AI reasoning block, folded by default. The title shimmers while the model is thinking, then reads Thought for Ns, with an animated height reveal.npx shadcn@latest add ui.minidev.pro/r/reasoning-block.jsonThinking BlockuiReact Thinking panel with a chevron header and aria-expanded that reveals the model's intermediate reasoning in muted text. Collapsed by default in chat.npx shadcn@latest add ui.minidev.pro/r/thinking-block.jsonTool Call CarduiAgent tool call card showing the tool name, argument preview, duration and a running, done or failed status chip, with collapsible monospace output below.npx shadcn@latest add ui.minidev.pro/r/tool-call-card.jsonSuggestion ChipsuiRow of rounded suggestion chips such as Summarize or Find bugs that call onSelect when clicked. Use them as prompt starters under a React AI chat input.npx shadcn@latest add ui.minidev.pro/r/suggestion-chips.json

Frequently asked questions

How do I stream a chat response in React without a provider SDK?

Call fetch, read res.body.getReader() in a loop, decode each chunk with TextDecoder using { stream: true }, and split complete lines out of a buffer. Update the assistant message in state as each event arrives.

How do I add a stop button to a streaming chat?

Create an AbortController per request, pass its signal to fetch, and call abort() from the stop button. The read loop throws, and you mark the partial message as stopped instead of failed by checking signal.aborted.

How should tool calls appear in a chat UI?

As compact cards inside the assistant turn that show the tool name, a short argument preview and a running, done or failed status, with the output collapsed. Upsert them by call id so the same card updates in place.

Can I build a ChatGPT style UI with shadcn components?

Yes. MiniDev UI is a shadcn-compatible registry, so npx shadcn@latest add https://ui.minidev.pro/r/chat-thread.json installs the thread, bubble and cursor into your project alongside any shadcn components you already use.

More guides