Skip to content
Scrim UI

AI Voice Assistant

A voice-first conversation — live waveform states, a recording input, a spoken transcript and a typed fallback.

voice-assistant.tsx267 lines5 componentsReact + Tailwind + shadcn/ui

npx shadcn@latest add @scrimui/voice-assistant
Agent promptClaude Code · Cursor · any agent
Add the AI Voice Assistant pattern from the Scrim UI registry to this project — a complete screen, not a single component.

AI Voice Assistant — A voice-first conversation — live waveform states, a recording input, a spoken transcript and a typed fallback.

## 1. Install

```bash
npx shadcn@latest add @scrimui/voice-assistant
```

This writes the screen to `components/blocks/voice-assistant.tsx` and pulls in the components it is built from, each landing at `components/ui/`. Everything is plain React + Tailwind with no runtime dependencies. The imports in the block already point at those paths, so it compiles as installed.

## 2. What it is made of

- `voice-input` — Voice Input
- `voice-waveform` — Voice Waveform
- `voice-conversation` — Voice Conversation
- `streaming-message` — Streaming Message
- `prompt-input` — Prompt Input

Each is a separate file you can edit or replace on its own; the block is the arrangement, not a monolith.

## 3. Rules this layout depends on

Keep these when adapting the screen — they are the reasons it works, and they are easy to break while restyling:

- Mirror the audio state in the UI — listening, recording and speaking should each have a distinct visual, never just a spinner.
- Keep the transcript readable: committed turns stay as text while the live waveform reflects only the current utterance.
- Offer a typed fallback under the mic — voice-first does not mean voice-only, and transcription errors need an escape hatch.
- Stream the spoken answer as text too, so the reply stays reviewable and searchable after the audio finishes.
- Stop controls must end the turn cleanly — stopping the reply, the recording or the utterance should never orphan the transcript.

## 4. Do not do these

- Showing a static mic icon with no state change — the user cannot tell if the assistant is actually listening.
- Letting the waveform animate when nothing is happening; an idle waveform is noise, not information.
- No way to cancel a recording, forcing the user to finish a sentence they never started.
- Voice-only with no text transcript — anything the user can hear should be readable afterward.

The demo content in the file — messages, file names, model names — is placeholder. Replace it with this project's real data and wire the handlers to real state rather than shipping the stubs.

Reference: https://scrimui.dev/patterns/voice-assistant

Installs the screen and every component it is built from, and carries the layout rules from this page so an agent does not restyle them away.

Live Preview

Voice Assistant

Hands-free answers

Idle

Tap the mic to start talking

AI
AssistantNow

Hi, I’m your voice assistant. Tap the mic and talk, or type below — I’ll answer out loud.

Click to talk

Built from these components

Pattern code

This file composes the components above. Copy each component from its page, then this pattern file wires them together.

voice-assistant.tsx
"use client";

import { Badge } from "@/components/ui/badge";
import { Card } from "@/components/ui/card";
import * as React from "react";
import { VoiceInput } from "../../voice-input/voice-input";
import { VoiceWaveform, type WaveformState } from "../../voice-waveform/voice-waveform";
import { VoiceConversation, type VoiceTurn } from "../../voice-conversation/voice-conversation";
import { StreamingMessage } from "../../streaming-message/streaming-message";
import { PromptInput } from "../../prompt-input/prompt-input";

/* ------------------------------------------------------------------ */
/* Copy / data                                                         */
/* ------------------------------------------------------------------ */

type Stage = "idle" | "listening" | "recording";

const SPOKEN =
  "What does streaming versus waiting for the full reply do for perceived latency?";
const SPOKEN_WORDS = SPOKEN.split(" ");

const REPLIES = [
  "Streaming lands the first token in milliseconds, which makes a reply feel instant. Keep the reveal smooth, offer a stop control, and only surface citations once the claim is actually grounded.",
  "A voice-first interface should mirror its state out loud: listening, recording, speaking. The waveform is the visual echo of what the assistant is doing right now.",
];

const STAGE_TEXT: Record<Stage, string> = {
  idle: "Tap the mic to start talking",
  listening: "Listening…",
  recording: "Recording — tap stop when you are done",
};

const SPEAKING_TEXT = "Speaking…";

/* ------------------------------------------------------------------ */
/* Icons                                                               */
/* ------------------------------------------------------------------ */

function MicIcon() {
  return (
    <svg
      viewBox="0 0 24 24"
      fill="none"
      stroke="currentColor"
      strokeWidth="1.5"
      strokeLinecap="round"
      strokeLinejoin="round"
      width="15"
      height="15"
    >
      <path d="M12 2a3 3 0 0 0-3 3v7a3 3 0 0 0 6 0V5a3 3 0 0 0-3-3z" />
      <path d="M19 10v2a7 7 0 0 1-14 0v-2" />
      <path d="M12 19v3" />
    </svg>
  );
}

/* ------------------------------------------------------------------ */
/* Helpers                                                             */
/* ------------------------------------------------------------------ */

function waveState(stage: Stage, streaming: boolean): WaveformState {
  if (streaming) return "speaking";
  if (stage === "recording") return "recording";
  if (stage === "listening") return "listening";
  return "idle";
}

function stageText(stage: Stage, streaming: boolean) {
  return streaming ? SPEAKING_TEXT : STAGE_TEXT[stage];
}

function StatusChip({ stage, streaming }: { stage: Stage; streaming: boolean }) {
  let label = "Idle";

  let cls = "bg-muted text-muted-foreground";
  if (streaming) {
    label = "Speaking";
    cls = "bg-muted text-foreground";
  } else if (stage === "recording") {
    label = "Recording";
    cls = "bg-red-100 text-red-700 dark:bg-red-900/40 dark:text-red-300";
  } else if (stage === "listening") {
    label = "Listening";
    cls = "bg-muted text-foreground";
  }
  return (
    <Badge variant="secondary" className={`rounded-full px-2 py-0.5 text-xs font-medium ${cls}`}>{label}</Badge>
  );
}

/* ------------------------------------------------------------------ */
/* VoiceAssistantPattern                                               */
/* ------------------------------------------------------------------ */

export function VoiceAssistantPattern() {
  const [turns, setTurns] = React.useState<VoiceTurn[]>([
    {
      id: "1",
      role: "assistant",
      text: "Hi, I’m your voice assistant. Tap the mic and talk, or type below — I’ll answer out loud.",
      time: "Now",
    },
  ]);
  const [stage, setStage] = React.useState<Stage>("idle");
  const [transcript, setTranscript] = React.useState("");
  const [reply, setReply] = React.useState<string | null>(null);
  const [streaming, setStreaming] = React.useState(false);

  const scrollRef = React.useRef<HTMLDivElement>(null);
  const idRef = React.useRef(2);
  const replyRef = React.useRef(0);
  const wordRef = React.useRef(0);
  const typingRef = React.useRef<number | null>(null);

  React.useEffect(() => {
    const el = scrollRef.current;
    if (el) el.scrollTop = el.scrollHeight;
  }, [turns, reply, stage]);

  const recordingTime = `0:0${Math.min(3 + Math.ceil(transcript.length / 14), 9)}`;

  function clearTyping() {
    if (typingRef.current !== null) {
      window.clearInterval(typingRef.current);
      typingRef.current = null;
    }
  }

  function startListening() {
    if (streaming) {
      setReply(null);
      setStreaming(false);
    }
    clearTyping();
    setTranscript("");
    setStage("listening");
    window.setTimeout(() => {
      setStage("recording");
      wordRef.current = 0;
      typingRef.current = window.setInterval(() => {
        wordRef.current += 1;
        setTranscript(SPOKEN_WORDS.slice(0, wordRef.current).join(" "));
        if (wordRef.current >= SPOKEN_WORDS.length) clearTyping();
      }, 230);
    }, 900);
  }

  function cancelRecording() {
    clearTyping();
    setTranscript("");
    setStage("idle");
  }

  function stopRecording() {
    clearTyping();
    const text = transcript.trim();
    setTranscript("");
    if (!text) {
      setStage("idle");
      return;
    }
    setStage("idle");
    setTurns((t) => [...t, { id: String(idRef.current++), role: "user", text, time: recordingTime }]);
    beginReply();
  }

  function beginReply() {
    const text = REPLIES[replyRef.current % REPLIES.length];
    replyRef.current += 1;
    setReply(text);
    setStreaming(true);
  }

  function finishReply() {
    if (reply) {
      setTurns((t) => [
        ...t,
        { id: String(idRef.current++), role: "assistant", text: reply, time: "Now" },
      ]);
    }
    setReply(null);
    setStreaming(false);
  }

  function stopReply() {
    setReply(null);
    setStreaming(false);
  }

  function submitText(value: string) {
    setTurns((t) => [...t, { id: String(idRef.current++), role: "user", text: value, time: "Now" }]);
    beginReply();
  }

  return (
    <Card className="gap-0 py-0 flex h-[560px] flex-col overflow-hidden rounded-xl border border-border bg-card">
      {/* Header */}
      <div className="flex items-center justify-between border-b border-border px-4 py-3">
        <div className="flex items-center gap-2.5">
          <span className="flex h-8 w-8 items-center justify-center rounded-lg bg-muted text-muted-foreground">
            <MicIcon />
          </span>
          <div>
            <p className="text-sm font-semibold text-foreground">
              Voice Assistant
            </p>
            <p className="text-xs text-muted-foreground">Hands-free answers</p>
          </div>
        </div>
        <StatusChip stage={stage} streaming={streaming} />
      </div>

      {/* Live waveform strip */}
      <div className="flex items-center gap-3 border-b border-border px-4 py-2.5">
        <VoiceWaveform
          state={waveState(stage, streaming)}
          bars={22}
          className="h-7 w-40 shrink-0 text-muted-foreground"
        />
        <p className="truncate text-xs text-muted-foreground">{stageText(stage, streaming)}</p>
      </div>

      {/* Messages */}
      <div ref={scrollRef} className="flex-1 space-y-5 overflow-y-auto px-4 py-5 sm:px-6">
        <VoiceConversation turns={turns} />

        {reply && (
          <div className="flex items-start gap-3">
            <Badge variant="secondary" className="mt-3 shrink-0 rounded-full border border-border bg-muted px-2.5 py-1 text-xs font-medium text-muted-foreground">
              Speaking
            </Badge>
            <div className="min-w-0 flex-1">
              <StreamingMessage
                text={reply}
                isStreaming={streaming}
                speed={2}
                showActions={false}
                onStop={stopReply}
                onComplete={finishReply}
              />
            </div>
          </div>
        )}
      </div>

      {/* Controls */}
      <div className="space-y-2 border-t border-border px-4 py-3">
        <VoiceInput
          state={stage === "recording" ? "recording" : "idle"}
          recordingTime={recordingTime}
          transcript={transcript}
          onStart={startListening}
          onStop={stopRecording}
          onCancel={cancelRecording}
        />
        <PromptInput
          placeholder="Type instead…"
          onSubmit={submitText}
          showWebSearch={false}
          showTools={false}
          disabled={streaming}
        />
      </div>
    </Card>
  );
}

When to use it

  • Mirror the audio state in the UI — listening, recording and speaking should each have a distinct visual, never just a spinner.
  • Keep the transcript readable: committed turns stay as text while the live waveform reflects only the current utterance.
  • Offer a typed fallback under the mic — voice-first does not mean voice-only, and transcription errors need an escape hatch.
  • Stream the spoken answer as text too, so the reply stays reviewable and searchable after the audio finishes.
  • Stop controls must end the turn cleanly — stopping the reply, the recording or the utterance should never orphan the transcript.

What breaks in production

  • Showing a static mic icon with no state change — the user cannot tell if the assistant is actually listening.
  • Letting the waveform animate when nothing is happening; an idle waveform is noise, not information.
  • No way to cancel a recording, forcing the user to finish a sentence they never started.
  • Voice-only with no text transcript — anything the user can hear should be readable afterward.

More Patterns

AI Chat

The canonical chat interface — sidebar, streaming messages, prompt input with model selector, and sources.

AI Research Assistant

A research flow that shows search tool calls, reasoning, sources and a cited final answer.

AI Coding Agent

A coding run with agent status, tool calls, diffs and a human-in-the-loop approval gate.

Model & Memory Preferences

A preferences screen that picks the model, reasoning level and tools, and manages persistent memory.

Artifact Workspace

Chat on the left, generated output on the right — artifacts open from the answer, stream, version, and fail without breaking the conversation.

Document Q&A Workspace

Ask your own documents — upload and parse, cited answers with inspectable passages, an honest not-found state, and a visible context budget.

Structured Extraction & Review

Upload a document, watch fields fill in, then review the flagged ones — per-field confidence, corrections that keep the original, export earned.

Image Generation Studio

A one-screen generation studio — prompt composer and model picker beside a feed of queued, staged, blocked and ready image results with variants.

Multi-agent Ops Console

Watch a fleet of agents at once — parallel statuses, an inspectable handoff, a waiting approval, a failed child run, and per-run plus fleet cost.

Customer Support Copilot

Grounded reply drafts with citations, honest low-confidence answers, inline corrections, an approval gate on refunds, and a rating row on every draft.

Generative UI Dashboard

The model assembles a dashboard from a controlled widget registry — streamed props, an unsupported-request fallback, and widget clicks that re-enter the chat.