vox / Docs
llms.txt

Vox for Node

Transcribe, dictate and speak from a Bun or Node tool through Vox on your Mac.

@voxd/sdk lets a Bun or Node program use the speech engine in the Vox app on the same Mac. The model stays loaded across every tool that uses it.

1. Run Vox on your Mac

Install and open the Vox app, then check it: see Vox on your Mac.

bunx @voxd/cli@latest doctor        # expect ready: true

2. Install the SDK

bun add @voxd/sdk          # or: npm install @voxd/sdk

3. Transcribe a file and speak a line

import { writeFile } from "node:fs/promises";
import { VoxClient } from "@voxd/sdk";

const vox = new VoxClient({ clientId: "my-tool" });
await vox.connect();

await vox.preloadModel("parakeet:v3");
const result = await vox.transcribeFile("/absolute/path/to/audio.wav", "parakeet:v3");
console.log(result.text, `${Math.round(result.elapsedMs)} ms`);

const speech = await vox.synthesize("Hello from Vox.", { modelId: "avspeech:system", format: "wav" });
await writeFile("hello.wav", speech.audio);

vox.disconnect();
  • clientId names your tool in Vox’s timings. Pick something stable.
  • preloadModel loads the model before the first request. Skip it and the first transcription waits for the load.
  • File paths must be absolute: Vox reads the file itself.
  • avspeech:system is the Mac’s system voice. For gpt-4o-mini-tts, add an OpenAI key in the Vox app, or pass credentials: { OPENAI_API_KEY } in the options.

4. Dictate with live text

Vox records from the Mac’s microphone and streams partial text while the user speaks:

const session = vox.createLiveSession();
session.on("partial", ({ text }) => process.stdout.write(`\r${text}`));

const done = session.start();
setTimeout(() => session.stop(), 5000);
console.log("\n" + (await done).text);

The first live session makes macOS ask whether Vox may use the microphone.

Runnable example

examples/node-hello:

cd examples/node-hello
bun install
bun run index.ts path/to/audio.wav

Next

Search

Find docs fast