Transcription models ​
Transcription models convert audio bytes into normalized text. They fit upload processors, meeting notes, support-call review, accessibility, and media pipelines.
1. Create a transcription model ​
Construct the provider model in server-side configuration.
import { OpenAIClient } from '@anvia/openai'
const client = new OpenAIClient({ apiKey })
export const transcriptionModel = client.transcriptionModel({ modelId: 'whisper-1' })OpenAI, Gemini, and Grok provide v1 transcription adapters. Supported formats, size limits, languages, and optional parameters vary by provider.
2. Read the audio bytes ​
Validate authorization, media type, and size before loading an upload into memory.
import { readFile } from 'node:fs/promises'
const audioPath = 'support-call.wav'
const audio = await readFile(audioPath)transcribe() accepts audio bytes as audio.data, a Uint8Array or ArrayBuffer in its options object. Empty audio is rejected before the provider is called.
3. Transcribe the audio ​
Pass one options object containing audio: { data, filename, mediaType? } and model.
import { transcribe } from '@anvia/core/transcription'
import { transcriptionModel } from './models'
const transcript = await transcribe({
audio: {
data: audio,
filename: audioPath,
mediaType: 'audio/wav'
},
model: transcriptionModel,
language: 'en',
prompt: 'Transcribe the customer support call exactly.',
temperature: 0
})
console.log(transcript.text)A useful filename helps the provider infer the media type; pass an explicit mediaType on the audio object when the filename is ambiguous. Language, prompt, temperature, provider options, retries, and abortSignal are optional.
4. Handle the result safely ​
The normalized result contains text and the raw provider response.
await transcripts.save({
recordingId,
text: transcript.text,
})Audio and transcripts may contain private or regulated information. Define access, retention, deletion, redaction, and downstream-use rules before sending transcripts to agents, indexes, analytics, or logs.
For document and image extraction, continue with OCR models.