The voice-first API for real-time avatars.

Orvyn lets developers put a rendered talking avatar into any product: create a session over REST, stream microphone audio over WebSocket, and receive the avatar as a single fMP4 stream.

REST session in → WebSocket mic up → fMP4 avatar out

ORVYN AVATAR DEMO
Sign in to talk — opens in your console
Arnold — Orvyn starter avatar

Talk to Arnold and the other starter avatars — free in your console.

Connection Status Notice
Unable to connect to Orvyn demo environment.

Three calls from mic to face.

The whole pipeline is three steps. Your backend creates the session, the browser speaks into a WebSocket, and the avatar comes back rendered.

01 // CREATE A SESSION

Create a session via REST

One authenticated call to POST /v1/sessions from your backend returns the session plus a short-lived client_token — delivered exactly once, five-minute TTL, refreshable.

02 // STREAM MIC AUDIO

Stream mic audio over WebSocket

The browser opens WS /v1/sessions/{id}/ws and sends raw PCM16 microphone audio upstream. Speech events and turn boundaries stream back as JSON — nothing is recorded or stored.

03 // RENDERED SERVER-SIDE

Receive the fMP4 avatar stream

Neural rendering happens on Orvyn's servers. The browser receives a single fMP4 avatar stream and hands it to a plain <video> element. No codecs, no client GPU code.

The browser stays thin

Our rendering servers are the single source of truth for rendered audio and video. Your client never processes, replays, or buffers media — it speaks, listens, and displays. That is the whole client.

Everything a voice-first product needs.

No invented features here — these are the capabilities the Orvyn platform ships today, end to end.

REST session lifecycle

Create, inspect, and end sessions with a small, versioned REST API — full status visibility from creation to end.

WebSocket media

Raw PCM16 microphone audio upstream; JSON events and fMP4 media downstream over one authenticated socket.

Single-stream fMP4 rendering

The avatar arrives as one server-rendered fMP4 stream — lip-synced audio and video in a single feed, playable by any modern browser.

TypeScript SDK

A typed browser SDK for the full session lifecycle — connect, refresh, end — and a single fMP4 avatar stream into a standard <video> element.

Unified telemetry

PostHog product analytics and Sentry error tracking are wired through the platform, so every session is observable out of the box.

Clerk authentication

Sign-in and account management run through Clerk on the Orvyn dashboard — your API keys and sessions stay behind real auth.

A typed SDK, start to end.

The official TypeScript SDK covers the whole lifecycle: session creation, WebSocket media, token refresh, and server-rendered playback into a video element you own.

// npm install @avtr/web-sdk
import { AvtrWebSdk, createSession } from "@avtr/web-sdk";

// 1. Create a session (server-side, with your API key)
const session = await createSession(apiBaseUrl, apiKey, {});
//    → { sessionId, clientToken, expiresAt, … }

// 2. Connect: mic up (PCM16), avatar down (fMP4)
const sdk = new AvtrWebSdk({
  apiBaseUrl: "https://api.orvyn.live",
  videoElement: document.querySelector("video")!,
});
await sdk.connect(session.sessionId, session.clientToken);

sdk.on("speechStarted", () => console.log("user is speaking"));
// 3. Token near expiry: refresh and keep talking
await sdk.refreshNow();
// Done? End the session cleanly.
await sdk.endSession();

Give your product a voice. And a face.

Create your first session in minutes — REST in, a rendered talking avatar out.

Prefer to try before you sign up? Talk to a live avatar →