byteplus/seed-audio-1.0, served by Comfy Router from BytePlus.
Quick start
Create a key at platform.comfy.org/profile/api-keys and export it asCOMFY_API_KEY. The Python and TypeScript snippets use the Comfy SDKs (pip install comfy-sdk, npm install @comfyorg/sdk); the cURL snippet is the same call over raw HTTP.
Model ID: byteplus/seed-audio-1.0
Endpoint: POST https://api.comfy.org/v2/models/byteplus/seed-audio-1.0
Schema
Input
object
Output audio configuration.
boolean
Whether to enable the subtitle service. When enabled, the response includes utterance- and word-level timestamps (default: false).
string
Output audio format: wav (default), mp3, pcm or ogg_opus.Possible values:
wav, mp3, pcm, ogg_opusinteger
-50 to 100; 100 means 2.0x volume, -50 means 0.5x volume (default: 0).Range:
-50 to 100integer
-12 to 12 (default: 0).Range:
-12 to 12integer
Output sample rate: 8000, 16000, 24000 (default), 32000, 44100 or 48000.Possible values:
8000, 16000, 24000, 32000, 44100, 48000integer
-50 to 100; 100 means 2.0x speed, -50 means 0.5x speed (default: 0).Range:
-50 to 100string
Model identifier. Supported models: seed-audio-1.0 and seed-audio-1.0-multilingual. A direct v1 call to POST /proxy/byteplus/api/v3/tts/create MUST supply it — the proxy refuses any other value, and an omitted one, with a 400. It is NOT in this schema’s
required list because Comfy Router fills it from the {model} path segment of /v2/models/byteplus/{model}, so a Router caller omits it.object[]
Reference resources. Omit for text-only generation. Up to 3 audio references (each up to 30 seconds and 10 MB decoded; wav, mp3, pcm or ogg_opus) or exactly 1 image reference (up to 10 MB decoded; jpeg, png or webp). Image references cannot be mixed with audio references. Those ceilings are BytePlus’s and are stated in DECODED bytes: base64 inflates an asset by about 4/3, and a Comfy Router call is capped at 10 MiB for the whole JSON document, so an inline
audio_data/image_data reference sent to /v2/models/byteplus/{model} has to stay under roughly 7.5 MB decoded. Use audio_url/image_url for anything larger — a reference at BytePlus’s own 10 MB maximum is refused 413 before validation runs if it is inlined.string
Base64-encoded reference audio.
string
URL of a remote reference audio file.
string
Base64-encoded reference image.
string
URL of a remote reference image.
string
Voice ID. Can be a supported Doubao TTS voice or a voice-clone voice ID.
string
required
Prompt or text to synthesize (1 to 3,000 characters; an empty string is refused). When audio references are provided, reference them by order using @Audio1, @Audio2 and @Audio3.
object
Watermark configuration object. An empty object is accepted.
object
Implicit watermark. Adds metadata to the synthesized audio header.
string
Name or code of the synthesis service provider.
string
Name or code of the content distribution service provider.
boolean
Whether to enable the implicit watermark (default: false).
string
Content production ID.
string
Content distribution ID.
boolean
Explicit watermark switch. Adds an audio rhythm marker at the end of the synthesized audio (default: false).
GET /v2/models/byteplus/seed-audio-1.0/openapi.json, the same document it validates a call against before the request reaches the provider.
Output
string
Generated audio data, Base64-encoded.
integer
Status code. Refer to the official error-code document for details.
number
Duration after speed or post-processing, in seconds.Format:
doublestring
Status details.
number
Original model output duration in seconds. Used for billing and capped at 120 seconds.Format:
doubleobject
Subtitle information for the audio. Present only when audio_config.enable_subtitle is set to true in the request.
object[]
Utterance-level subtitle information.
integer
Segment end time, in milliseconds from the beginning of the audio.
integer
Segment start time, in milliseconds from the beginning of the audio.
string
The complete text of the segment.
string
Subtitle text corresponding to the audio.
object[]
Word-level subtitle information.
integer
Segment end time, in milliseconds from the beginning of the audio.
integer
Segment start time, in milliseconds from the beginning of the audio.
string
The complete text of the segment.
string
Temporary audio URL, valid for 2 hours.
Examples
Input
Output
Before you ship
The snippets above are the shortest working call. Three things are the same for every model and are documented once on the Comfy Router headers page: send anIdempotency-Key on every paid call and reuse it when you retry, expect the connection to be held up to Router’s 10 minute deadline, and keep X-Comfy-Request-Id from every response. The SDKs do all three for you; the cURL tab does none of them. On failure, X-Comfy-Error-Type names the bucket, and a 422 means the body failed the model’s schema and was never billed.
Headers
Authentication, idempotency, request IDs, error buckets, retry pacing, spend limits.
Quick Start
Typed error handling in Python and TypeScript, reading the 422, walking the catalog.
Limitations
What Router does not do today, and what to use instead.