The voice-first API for real-time avatars.
Orvyn lets developers put a rendered talking avatar into any product: create a session over REST, stream microphone audio over WebSocket, and receive the avatar as a single fMP4 stream.
REST session in → WebSocket mic up → fMP4 avatar out
Talk to Arnold and the other starter avatars — free in your console.
Three calls from mic to face.
The whole pipeline is three steps. Your backend creates the session, the browser speaks into a WebSocket, and the avatar comes back rendered.
Create a session via REST
One authenticated call to POST /v1/sessions from your
backend returns the session plus a short-lived
client_token — delivered exactly once, five-minute TTL,
refreshable.
Stream mic audio over WebSocket
The browser opens WS /v1/sessions/{id}/ws and sends raw
PCM16 microphone audio upstream. Speech events and turn boundaries
stream back as JSON — nothing is recorded or stored.
Receive the fMP4 avatar stream
Neural rendering happens on Orvyn's servers. The browser receives a
single fMP4 avatar stream and hands it to a plain
<video> element. No codecs, no client GPU code.
The browser stays thin
Our rendering servers are the single source of truth for rendered audio and video. Your client never processes, replays, or buffers media — it speaks, listens, and displays. That is the whole client.
Everything a voice-first product needs.
No invented features here — these are the capabilities the Orvyn platform ships today, end to end.
REST session lifecycle
Create, inspect, and end sessions with a small, versioned REST API — full status visibility from creation to end.
WebSocket media
Raw PCM16 microphone audio upstream; JSON events and fMP4 media downstream over one authenticated socket.
Single-stream fMP4 rendering
The avatar arrives as one server-rendered fMP4 stream — lip-synced audio and video in a single feed, playable by any modern browser.
TypeScript SDK
A typed browser SDK for the full session lifecycle — connect, refresh, end — and a single fMP4 avatar stream into a standard <video> element.
Unified telemetry
PostHog product analytics and Sentry error tracking are wired through the platform, so every session is observable out of the box.
Clerk authentication
Sign-in and account management run through Clerk on the Orvyn dashboard — your API keys and sessions stay behind real auth.
A typed SDK, start to end.
The official TypeScript SDK covers the whole lifecycle: session creation, WebSocket media, token refresh, and server-rendered playback into a video element you own.
// npm install @avtr/web-sdk
import { AvtrWebSdk, createSession } from "@avtr/web-sdk";
// 1. Create a session (server-side, with your API key)
const session = await createSession(apiBaseUrl, apiKey, {});
// → { sessionId, clientToken, expiresAt, … }
// 2. Connect: mic up (PCM16), avatar down (fMP4)
const sdk = new AvtrWebSdk({
apiBaseUrl: "https://api.orvyn.live",
videoElement: document.querySelector("video")!,
});
await sdk.connect(session.sessionId, session.clientToken);
sdk.on("speechStarted", () => console.log("user is speaking"));
// 3. Token near expiry: refresh and keep talking
await sdk.refreshNow();
// Done? End the session cleanly.
await sdk.endSession();