Skip to content

OpenAI

OpenAI ships sync transcription via REST, realtime transcription via WebSocket, and text-to-speech via chunked HTTP. Each lives at its own subpath so the realtime peer dep doesn’t infect sync-only builds.

Install

Terminal window
pnpm add @effect-uai/core @effect-uai/openai effect

The realtime transcriber additionally needs ws (a peer dep):

Terminal window
pnpm add ws

ws is only pulled in by @effect-uai/openai/OpenAIRealtimeTranscriber. The sync OpenAITranscriber and OpenAISynthesizer paths don’t require it: edge / browser builds stay slim.

Layers

LayerRegistersCapability markers
@effect-uai/openai/OpenAITranscriberOpenAITranscriber + Transcribernone
@effect-uai/openai/OpenAIRealtimeTranscriberOpenAIRealtimeTranscriber + TranscriberSttStreaming
@effect-uai/openai/OpenAISynthesizerOpenAISynthesizer + SpeechSynthesizernone (no TtsIncrementalText: OpenAI has no /stream-input endpoint)
import { Config, Effect, Layer } from "effect"
import { FetchHttpClient } from "effect/http"
import { layer as transcriberLayer } from "@effect-uai/openai/OpenAITranscriber"
import { layer as realtimeLayer } from "@effect-uai/openai/OpenAIRealtimeTranscriber"
import { layer as synthLayer } from "@effect-uai/openai/OpenAISynthesizer"
const openai = Layer.unwrap(
Effect.gen(function* () {
const apiKey = yield* Config.Redacted("OPENAI_API_KEY")
return Layer.mergeAll(
transcriberLayer({ apiKey }), // sync STT
realtimeLayer({ apiKey }), // streaming STT
synthLayer({ apiKey }), // sync + chunked TTS
)
}),
)
const mainLayer = openai.pipe(Layer.provide(FetchHttpClient.layer))

Models

STT

ModelSyncStreamingNotes
gpt-transcribe✓✓Current sync model; plain text only
gpt-live-transcribenone✓Current streaming model; plain text only
gpt-realtime-whispernone✓Streaming Whisper
gpt-4o-transcribe✓✓Deprecated, shutdown 2027-02-26
gpt-4o-mini-transcribe✓✓Deprecated, shutdown 2027-02-26
whisper-1✓noneDeprecated, shutdown 2027-02-26; only model with wordTimestamps

wordTimestamps: true requires whisper-1. Passing it to another model surfaces the provider’s wire rejection (HTTP 400) rather than a pre-send error. diarization is narrowed off OpenAITranscribeRequest (OpenAI’s transcription endpoint has none).

TTS

ModelStreamingNotes
gpt-4o-mini-ttschunked HTTPCurrent steerable model; honors instructions
tts-1 / tts-1-hdchunked HTTPLegacy; ignore instructions silently

Stock voices (no custom-voice path): alloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verse. ballad, coral, and verse are gpt-4o-mini-tts-only. Because there’s no clone path, OpenAISynthesizeRequest.voiceId narrows to the stock-only literal union. Passing an arbitrary string is a type error.

Request shape

// STT sync
type OpenAITranscribeRequest = {
readonly model: OpenAITranscribeModel
readonly audio: AudioSource
readonly language?: string
readonly prompt?: string // free-form prose context, mapped to OpenAI's prompt
readonly biasingTerms?: ReadonlyArray<string> // warnDropped (no keyterm field)
readonly wordTimestamps?: boolean // whisper-1 only
readonly temperature?: number
readonly fileName?: string // overrides multipart filename
}
// TTS sync + chunked
type OpenAISynthesizeRequest = {
readonly model: OpenAITtsModel
readonly voiceId: OpenAIVoiceId // stock-only literal union
readonly text: string
readonly outputFormat?: AudioFormat
readonly speed?: number
readonly instructions?: string // gpt-4o-mini-tts only
}

instructions is a free-form prompt for tone, emotion, pacing: “sound apologetic,” “read this slowly with emphasis on the second sentence.” Honored only by gpt-4o-mini-tts; silently ignored by the legacy tts-1 family.

Wire / auth notes

Realtime STT opens wss://api.openai.com/v1/realtime?intent=transcription and sends a session.update with type: "transcription" (the GA wire; the old OpenAI-Beta: realtime=v1 header was shut down 2026-05-12). The upgrade needs an Authorization: Bearer … header, which browser WebSocket can’t set, so OpenAIRealtimeTranscriber uses the ws peer dep to construct the socket. That’s why this transcriber lives at a separate subpath. Use it from Node / Bun; for browser deployments, proxy through a server.

Realtime expects PCM s16le at 24 kHz (not 16 like most other providers). Set inputFormat accordingly on the streaming request, or the upstream rejects the audio.

vadEvents maps to audio.input.turn_detection and is on unless you pass false. It is what makes the server commit each turn, and a committed turn is what produces the final transcript alongside the speech-started / utterance-ended boundaries. A model that transcribes continuously instead of segmenting rejects the field with invalid_value; pass vadEvents: false there and expect partial events only, since nothing commits the audio buffer.

Output formats for TTS: mp3, opus, aac, flac, wav, pcm. pcm is 24 kHz mono s16le, suitable for direct AudioWorklet playback.

Errors

Standard HTTP → AiError mapping applies:

StatusError
429AiError.RateLimited
408/504AiError.Timeout
401AiError.AuthFailed (auth)
>= 500AiError.Unavailable
other 4xxAiError.InvalidRequest

wordTimestamps: true against a non-whisper-1 model → AiError.Unsupported at request time.

See also

  • Speech overview: generic tags and capability markers.
  • Voice loop: uses ElevenLabs by default; the recipe’s runPipeline typechecks against either provider via the marker contract.
  • Streaming transcription: default provider is OpenAI Realtime.