Authentication
For authenticated endpoints, send your API key in `Authorization: Bearer <key>` or `x-api-key`. API keys are shown when created and can be accessed from your account management page.
Static-hosted account management is available at `/dashboard` using an API key as your primary credential.
Bootstrap without checkout is also available using a Google ID token on `/api/v1/auth/google` or `/api/v1/keys`.
Email one-time-code bootstrap is also available using `/api/v1/auth/email/start` and `/api/v1/auth/email/verify`.
Hosted browser flows can set a signed browser session cookie via `/api/v1/auth/browser/google` or `/api/v1/auth/browser/email/verify`, then inspect or clear it with `/api/v1/auth/browser/session`.
For local coding agents, device-link bootstrap is available using `/api/v1/device/link/start`, `/api/v1/device/link/approve`, and `/api/v1/device/link/poll`.
Google auth key mint request
curl -X POST "$API_BASE/api/v1/auth/google" \
-H "Content-Type: application/json" \
-d '{
"idToken": "eyJhbGciOiJSUzI1NiIs...",
"keyLabel": "Google bootstrap key"
}'Hosted browser Google auth request
curl -X POST "$API_BASE/api/v1/auth/browser/google" \
-H "Content-Type: application/json" \
-d '{
"idToken": "eyJhbGciOiJSUzI1NiIs..."
}'Bootstrap key mint via /api/v1/keys
curl -X POST "$API_BASE/api/v1/keys" \
-H "Content-Type: application/json" \
-d '{
"idToken": "eyJhbGciOiJSUzI1NiIs...",
"label": "Google bootstrap key"
}'Inspect browser session
curl -X GET "$API_BASE/api/v1/auth/browser/session" \
-H "Cookie: killa_tamata_browser_session=<session-cookie>"
Start local device-link request
curl -X POST "$API_BASE/api/v1/device/link/start" \
-H "Content-Type: application/json" \
-d '{
"keyLabel": "Default API Key"
}'Qwen 3.8 27B multimodal inference
Use the OpenAI-compatible /api/v1 base with model qwen3.8-27b-uncensored. The endpoint supports text, separate reasoning_content, SSE streaming, and modern function tools with parallel calls. It is intentionally not part of Studio.
All GPU-ready work shares a FIFO queue. A non-streaming completion may return 202 Accepted with a gateway-local Location: /api/v1/jobs/{id}. Follow that job instead of resubmitting it, poll with the state-aware Retry-After, then fetch its /result. Submit directly to POST /api/v1/jobs when durable waiting is preferred.
Images and limits
- Use ordered text and image_url content blocks with base64 data URIs only.
- PNG, JPEG, WebP, and GIF are accepted; remote image URLs are rejected.
- Maximum 8 images and 10 MiB decoded per image.
- Convex limits each gateway request and response to 20 MiB, stricter than the upstream 32 MiB body limit. Oversized JSON returns 502; an already-started SSE response is interrupted if it crosses the limit.
Billing and safety
- $0.45 per million prompt tokens and $2.50 per million completion tokens; reasoning is completion usage.
- A $0.32768 hold is reserved first, then unused funds are refunded from authoritative usage.
- Definitive upstream failures or cancellation before output refund the hold; expired results and ambiguous response loss retain it.
- The gateway stores ownership/billing/job metadata, never prompts, images, reasoning, tool arguments/results, or generated content.
Prefix caching
- Set a stable
prompt_cache_key, up to 128 characters, when requests reuse the same long input prefix. - Keys are account-scoped, replaced with opaque upstream handles, and never persisted. Do not include secrets or personal data.
- Keyed responses may report best-effort
usage.prompt_tokens_details.cached_tokens telemetry. - Caching is an optimization, not a guarantee or billing discount; full
prompt_tokens remain billable.
This is an uncensored checkpoint. Apply the moderation, age-gating, output review, and domain-specific safety controls your application requires. Generation submissions are never retried automatically. Preserve one stable idempotency key across transport retries and honor Retry-Afterfor async polling and retryable streaming 429 or service 503 responses.
Text and reasoning request
curl -X POST "$API_BASE/api/v1/chat/completions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: qwen-text-001" \
-d '{
"model": "qwen3.8-27b-uncensored",
"prompt_cache_key": "support-agent-v3:conversation-8421",
"reasoning_effort": "medium",
"messages": [{"role": "user", "content": "Explain the image-analysis process concisely."}],
"max_completion_tokens": 256
}'Durable job submit, status, and result
curl -i -X POST "$API_BASE/api/v1/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: qwen-job-001" \
-d '{
"model": "qwen3.8-27b-uncensored",
"messages": [{"role": "user", "content": "Write a concise product description."}]
}'
# Honor Retry-After while the job is nonterminal.
curl -i "$API_BASE/api/v1/jobs/$JOB_ID" \
-H "Authorization: Bearer $API_KEY"
# Returns 202 while pending, or the OpenAI completion at 200.
curl -i "$API_BASE/api/v1/jobs/$JOB_ID/result" \
-H "Authorization: Bearer $API_KEY"Inline image-analysis request
curl -X POST "$API_BASE/api/v1/chat/completions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b-uncensored",
"reasoning_effort": "none",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Describe the image and transcribe visible text."},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,/9j/4AAQSk..."}}
]
}]
}'Required parallel function-tool loop
# 1) Send tools with tool_choice "required" (or "auto").
{
"model": "qwen3.8-27b-uncensored",
"messages": [{"role":"user","content":"Compare weather in Paris and Tokyo."}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Read current weather",
"parameters": {
"type": "object",
"properties": {"city":{"type":"string"}},
"required": ["city"]
}
}
}],
"tool_choice": "required",
"parallel_tool_calls": true
}
# 2) Validate and execute every returned tool call, then continue with:
[
{"role":"assistant","content":null,"reasoning_content":"...","tool_calls":[
{"id":"call_paris","type":"function","function":{"name":"get_weather","arguments":"{"city":"Paris"}"}},
{"id":"call_tokyo","type":"function","function":{"name":"get_weather","arguments":"{"city":"Tokyo"}"}}
]},
{"role":"tool","tool_call_id":"call_paris","content":"{"temperature_c":20}"},
{"role":"tool","tool_call_id":"call_tokyo","content":"{"temperature_c":27}"}
]
# Preserve the assistant message, match every tool_call_id, and send the same tools again.SSE streaming request with usage
curl -N -X POST "$API_BASE/api/v1/chat/completions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b-uncensored",
"messages": [{"role": "user", "content": "Give two concise observations."}],
"stream": true,
"stream_options": {"include_usage": true}
}'Balance and media usage
Submit jobs with task + input. Use this section as a quick reference for task behavior, billing expectations, and tuning controls.
Task quick map
12 production tasks
Choose one task per request and pair it with the matching input schema.
image.design
text-aware graphic design
image.edit
Qwen Image Edit
video.generate
H3 fast / balanced / quality
video.generate.reference
H3 Ref2VA Turbo + base reference
video.combine
clip stitch-down with audio preserved
video.resize
deterministic video rescale utility
audio.speak
Breeze, Gemini, and OmniVoice TTS
audio.annotation.reference
Whisper JSON, anchored or transcriptless
ace.step.create
Ace Step music
moss.sound.effect
MOSS sound effects
trellis.generate
Trellis 2 low-poly image-to-3D
Image and input handling
image.generate: set width + height directly when you need an explicit size override (optional, multiples of 32) up to 4MP.image.design: prompt-only graphics/posters with plain text or structured JSON prompt, preset quality or fast; output defaults to WebP quality 90 on a 1024x1024 canvas, and accepts explicit width + height only when provided together.image.edit: output resolution follows source metadata when available, else defaults to 1024x1024.outputFormat supports webp, png, and jpg; when omitted it defaults to webp.highDetail is an optional opt-in for image.generate and image.edit. It defaults to false, increases inference steps by 50%, and increases job cost by 50%. Only enable it when the caller explicitly wants a higher-detail image pass.- Source/reference images can be URL fields or inline base64 objects; inline payloads are staged on CDN (best-effort WebP conversion) and auto-cleaned.
- For both
image.generate and image.edit, use referenceImages (URL or inline entries). referenceImageUrls is still accepted as a URL-only legacy alias.
Audio usage and billing
- Set explicit durations:
ace.step.create 15..240s, moss.sound.effect 1..30s. - For polished long-form music, use
ace.step.create.input.qualityPreset=high_quality (108s, 192k). Internal XL Turbo sampler settings are fixed by the API. audio.speak supports three providers. Omit input.provider for OmniVoice compatibility, or set provider=gemini for direct Gemini 3.1 Flash TTS. Those two providers accept text 1..6000 chars.- Set
provider=breeze for Breeze-TTS-2 voice design, cloning, or voice direction. Design and direction require natural-language instructions; clone and direction require one reference audio source and its exact referenceTranscript. Use a reference you are authorized to use. Match instructions to the spoken language (English or Chinese). Output is mono 24 kHz Opus at 128k. Breeze accepts up to 1500 characters and 100 estimated speech seconds; for low-latency playback, keep coherent clips at or below 420 characters and submit later clips concurrently. Current optimized acceptance tests are English-only. Provisional pricing is $0.007 per estimated GPU second with a $0.07 minimum. The API returns the charged amount with each submission. - OmniVoice voice design still uses
voiceDescription 4..240 chars with comma-separated tags such as female, young adult, high pitch, american accent. Supported tags in this API release: gender female, male; age child, teenager, young adult, middle-aged, elderly; pitch very low pitch, low pitch, moderate pitch, high pitch, very high pitch; style whisper; accent american accent, british accent, australian accent, canadian accent, indian accent, chinese accent, korean accent, japanese accent, portuguese accent, russian accent. - Gemini requests use
voiceName and optional instructions instead of OmniVoice tags. Gemini v1 is single-speaker only, defaults to voiceName="Kore", and returns WAV output. - Gemini rejects
mode=voice_clone, referenceAudio*, referenceTranscript, voiceDescription, quality, and seed. Use OmniVoice when you need those controls. - For OmniVoice
audio.speak mode=voice_clone, use a matching transcript excerpt and keep the reference clip short. Start with 3..6s and add voiceDescription only when you need light delivery steering layered onto the cloned voice. OmniVoice accepts a fixed tag vocabulary rather than free-form prose. - OmniVoice upstream also documents Chinese-language attribute prompts, including Chinese dialect tags, but this API currently validates only the English tag set above on the OmniVoice path.
- Gemini billing now settles on completion after a submit-time hold. The final bill is
max(10_000, ceil(1.25 * (inputTextTokens * 1 + outputAudioTokens * 20))) usdMicros with outputAudioTokens = exactAudioSeconds * 25. Holds are based on exact Gemini prompt token counts plus a conservative output-duration estimate. - OmniVoice billing remains submit-time only. Voice clone currently uses the same billing baseline as voice design, and short clips still floor at
$0.02. audio.annotation.reference takes sourceAudioUrl/sourceAudio and optionally transcript, then returns a downloadable JSON annotation with lyrics, word timings, sections, beats, and QA metadata. If transcript is present, the gateway uses the reference alignment workflow. If it is omitted, the gateway switches to transcriptless ASR and defaults transcriptionModel to large-v3.- Ace/MOSS remain fixed at submit time. Audio annotation uses estimated duration from transcript length when present, and transcriptless ASR defaults to the
large-v3 estimate path. Gemini TTS is the only current audio flow that re-settles to an exact completion-time bill.
Video quality routing, multimodal H3 references, and reliability
video.generate supports image-to-video (startFrameImageUrl / startFrameImage, with sourceImageUrl / sourceImage accepted as aliases) and text-to-video (omit start/source image fields). Use quality=fast for the latest four-step FL2VA Turbo LoRA, quality=balanced (the default) for the latest eight-step Turbo LoRA, or quality=quality for base H3 at 20 steps. Deprecated low/regular aliases map to balanced; high maps to quality.- Optional
width + height are required together (multiples of 32) up to 4MP; otherwise use aspectRatio defaults. These values define the final output canvas (for example 1920x1088); do not post-rescale in clients unless you explicitly want a different deliverable size. - Optional final-frame steering via
finalFrameImageUrl or finalFrameImage. For all image fields, provide either a URL string or inline dataBase64 + mimeType (JPG/PNG/WEBP). Pricing is identical across image-to-video and text-to-video modes within a quality tier. Pricing scales from native frame count and final output resolution. The quality base-H3 tier adds the existing 50% surcharge. - Both video generation tasks accept one optional
watermark with an image plus integer x, y, width, and height in final-frame pixels. The complete box must fit on the delivered canvas. Compositing happens after resize, upscale, and RIFE; PNG/WebP alpha is preserved, while JPEG watermarks are opaque. A reference-job watermark does not count as a generation reference. - All H3 tiers generate video and audio together at a native
24fps on a canvas targeting a 768px short edge subject to the official 768x1344 pixel-area cap, applies RTXVideoSuperResolution at ULTRA quality only when the requested delivery canvas is larger, and uses RIFE to deliver 60fps by default. Fast and balanced use four and eight steps respectively with video/audio shifts 6/3; quality uses 20 steps without a Turbo LoRA. Fast and balanced currently have the same standard price; quality retains the 50% surcharge. extremeQuality=true remains a deprecated alias for quality. An aspect-only H3 request still uses the standard delivery canvas (for example 16:9 delivers1920x1088 from a 1344x768 native render); provide explicit width and height when the native-size deliverable is desired. video.generate.reference is always MiniMax H3 and accepts up to 9 images, 3 public MP4 videos (2-15 seconds each), and 3 standalone audio references. Optional startFrameImageUrl / startFrameImage andfinalFrameImageUrl / finalFrameImage anchor the opening and ending independently of the appearance slots. Guides use the centered native-canvas crop. A continuation inherits its opening and accepts an ending guide. Prompt tags follow array order: <Picture 1>, <Video 1>, and <Audio 1>. Reference videos guide visuals; add their soundtrack separately as a reference audio when it should also guide generated sound. Remote image and video references use provider-safe filenames; colliding image basenames are staged separately. H3 reads reference-video frames from the beginning, caps them to the output's native frame window, and may trim additional tail frames for frame-grid alignment. Fast uses the four-step Ref2VA Turbo v0.1 LoRA on its 544p profile; balanced uses the new eight-step Ref2VA Turbo v1.0 LoRA on its 768p profile. Both use shifts 12/3. Omitted or quality uses base H3 at 20 steps and 768p.video.generate supports optional start and final frames, or trim continuation with paired continuationSourceVideoUrl and continuationSeedTimeSeconds. Continuation probes the source, caches the immediate 56-frame (7/3-second) tail ending at the millisecond-quantized seed plus its terminal PNG, locks generated frame zero to that anchor, applies bounded RGB boundary grading, and conditions H3 Ref2VA on the tail and its embedded audio. Stitch the returned segment with video.combine using overlapFrames=1 and frameRate=60. Continuation cannot be combined with frame, overlap, or reference-audio inputs. Other generation requests reject overlap frames, exact reference audio, and firstPassImageStrength. Use multimodal reference generation whenever image, video, or audio conditioning is required.- Audio conditioning belongs on
video.generate.reference via referenceAudioUrls or referenceAudios. References guide H3's generated audio and motion; source audio is not copied verbatim. - Turbo graph details follow the upstream ComfyUI recipe:
LoraLoaderModelOnly at strength 1.2 for FL2VA Turbo (strength 1 for Ref2VA Turbo) feeds MiniMaxH3SigmaShift, which feeds both the guider and simple scheduler. Fast and balanced run exactly four and eight Euler steps; quality omits both nodes and runs the base graph for 20 steps. - H3 accepts 5-15 requested seconds at a fixed native 24 FPS. Exact H3 frame counts must satisfy
n % 17 = 5; common values are 124, 243, and 362. video.combine accepts legacy URL stitch-downs or structured trim, speed, color, and transition edits. Precision fields add curated output presets, fit/fill/custom framing, clip gain/mute/fades, linear or equal-power audio joins, and one looping soundtrack with bounded 0-24 dB ducking through FCSConcatVideosV4. Structured jobs preserve synchronized embedded audio, and fill silence for clips without audio. Authoritative probing happens before charging; edited output must be 1-600 seconds, and crossfades are capped to retain one 60 FPS frame per neighbor. Pricing is $0.01 base, $0.005 per input clip with a two-clip minimum, and $0.001 per output megapixel-second, plus $0.0005 per soundtrack output-second; structured duration uses probed metadata, trims, speed, and transition overlap.video.analyze returns authoritative source metadata, a CORS-verified 720p/30 FPS proxy with AAC audio, trim-ready WebP storyboard cues, mono waveform peak/RMS points, and loudness metadata. Canonical source analysis is idempotently cached per owner and all three artifacts are deleted 30 days after their last use. Pricing is $0.005 + $0.0005 per source megapixel-second. Optional frameTimestampsSeconds extracts at most eight original-resolution PNG stills at millisecond-quantized timestamps before the source end, for $0.001 per distinct frame. Timestamp selections are part of cache identity. Detailed job status returns their URLs, timestamps, and 30-day expiry in effectiveInput.extractedFrames.video.resize deterministically rescales an existing remote clip to explicit even width + height. Use frameRate to match the source clip FPS when you want exact timing preserved, and include sourceDurationSeconds when known so submit-time billing and ETA stay accurate. Current pricing is a utility formula: base $0.005 + megapixelSeconds * $0.001, where megapixelSeconds = MP * duration * (frameRate / 60).interpolationFps: set 0 to disable interpolation, or use 30..60 (default 60). MiniMax H3 uses native 24fps. Every path uses RIFE for the default 60fps delivery.
Image-to-video reliability warning
Image-to-video generation is probabilistic and can vary widely between runs. Some outputs will be frozen, warped, or otherwise unusable even with the same prompt and input image.
- Assume non-zero failure rate and run multiple candidates per request batch.
- Tune and select winners at low resolution first; upscale/extend only winning clips.
- Use the documented MiniMax fields below; keep one continuous Shot 1 unless a cut is genuinely required.
- Treat anchor-frame quality (mid-action, asymmetry, no text/signage) as the main quality lever.
MiniMax H3 prompt construction
For video.generate with a start frame, make the official I2VA alignment sentence the first line. In [Shot 1], explicitly carry forward the source subject's identity, appearance, colors, markings or clothing, composition, key objects, and spatial relationships before describing motion. The gateway adds this structure only when an I2V prompt is incomplete; a prompt that already contains the documented alignment sentence and all three core fields is passed through unchanged.
A start frame is cropped from the center to the video's native aspect ratio and is the exact opening frame. Cropping anchors frame zero; the prompt is what tells H3 which visible details must remain stable afterward. When text conflicts with the image, explicitly state whether the image or the requested transformation wins.
video.generate I2V prompt template
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic. The subject shown in <Picture 1> remains the same individual, preserving appearance, colors, markings, clothing, position, and scene layout. Describe the requested action and restrained camera motion here.
overall_soundscape: Describe ambience, physical sounds, and non-verbal sounds. Put dialogue in the shot description using <d>[Language] exact words</d>.
non_diegetic_music: N/A
video.generate.reference prompt template
subject_definitions:
<Subject 1> is the lead subject in <Picture 1>; preserve the subject's identity, appearance, colors, markings, proportions, and distinctive features.
<Video 1> provides camera movement and pacing.
<Audio 1> provides rhythm and generated-sound guidance.
summary: [reference generation + audio reference] Generate one continuous shot using <Subject 1>, the motion language of <Video 1>, and the timing of <Audio 1>.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - retain the referenced identity and appearance throughout.
<Video 1> (camera movement and pacing): reference - follow its motion without copying its visible subject.
<Audio 1>: reference - guide timing and sound without copying the signal verbatim.
detailed_description: Live-action, cinematic, with the visual style established by <Picture 1>. [Shot 1] <Subject 1> begins in a readable composition and performs the requested action while the camera follows <Video 1> with restrained movement. Motion and generated sound follow <Audio 1>.
overall_soundscape: Describe ambience, action sounds, and referenced audio behavior.
non_diegetic_music: N/A
- Full-reference prompts use six sections in order:
subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, and non_diegetic_music. - Define reusable people, animals, objects, scenes, or styles as
<Subject N> sourced from <Picture N>. Reserve picture labels for concrete frame/composition anchors and video labels for temporal structure, editing, or camera behavior. - In
retention_analysis, use an explicit relationship such as fully_preserved, partially_preserved, attribute_transfer, or weak_reference.
Turbo routing checklist
- Fast uses the latest 4-step Turbo LoRA; balanced uses the latest 8-step Turbo LoRA.
- Quality uses the base 20-step graph without Turbo and keeps the existing 50% surcharge.
- Video generation defaults to balanced; reference generation keeps its backward-compatible quality default.
- Legacy aliases remain accepted and are canonicalized before workflow, pricing, and status metadata.
- FL2VA Turbo uses LoRA strength 1.2; Ref2VA Turbo uses strength 1 and its task-specific 12/3 shifts.
- RIFE remains enabled by default for 60 FPS delivery; use interpolationFps 0 only for native 24 FPS.
Multimodal reference prompting
- Use Picture, Video, and Audio tags positionally in the same order as their request arrays.
- Prompt audio-conditioned shots as concrete performances, not generic portraits.
- Keep the subject readable whenever the face is foregrounded.
- Let mouth articulation, jaw travel, breath timing, shoulder rhythm, and phrase-timed gestures carry the sync.
- Favor restrained camera behavior. Prefer locked framing, a very gentle push, or a tiny lateral drift, with subject motion carrying the shot more than the camera.
- Start moving immediately and describe how the audio reference should affect performance and generated sound.
- Avoid prompts that mainly describe a seductive portrait, glamour still, or camera move. Those tend to animate the framing instead of the mouth.
Trellis 3D tuning notes
- Controls include
qualityPreset, targetFaceCount, mesh/texturing steps, texture size, and geometry cleanup knobs. - Low-poly mode also exposes
postprocessPositionEpsilon and postprocessNormalCreaseDeg for normal cleanup tuning. - Start with
qualityPreset=balanced. See /3d-models for benchmark-tuned examples.
Diagnostics and artifact window
POST /api/v1/media/jobs returns submit-time billing, hold/estimate details, effective input, and input-adjustment diagnostics. GET /api/v1/media/jobs?jobId=...&includeDetails=1 adds sanitized upstream result payloads plus the persisted effective input, adjustment trail, and any completion-time billing finalization metadata.- Set top-level
completionCallbackUrl when you want one terminal callback for completed, failed, terminated, and canceled/cancelled jobs. The POST body mirrors includeDetails=1 plus callbackType: "media_job_terminal". The callback must be a public HTTP(S) endpoint without URL credentials or redirects; private, loopback, and internal targets are rejected. - Download outputs within 48 hours. CDN cleanup runs after 72 hours, but availability past 48 hours is not guaranteed.
Audio pricing quick ref
audio.speak (Gemini)
max(10_000, ceil(1.25 * (inputTextTokens * 1 + outputAudioTokens * 20))) usdMicros. Finalized on completion with outputAudioTokens = seconds * 25.
audio.speak (OmniVoice)
max($0.01, estimatedGpuSeconds * $0.007). Estimated GPU seconds are derived from text length. Clone requests currently use the voice-design pricing baseline.
OmniVoice prompt tags
Supported English tags: female, male, child, teenager, young adult, middle-aged, elderly, very low pitch, low pitch, moderate pitch, high pitch, very high pitch, whisper, american accent, british accent, australian accent, canadian accent, indian accent, chinese accent, korean accent, japanese accent, portuguese accent, russian accent.
audio.annotation.reference
Estimated from transcript length when present: <=60s $0.03, <=120s $0.05, <=180s $0.07. Without a transcript the gateway defaults to transcriptless ASR with large-v3, which currently lands in the 180s band ($0.07).
ace.step.create
max($0.05, durationSeconds * qualityKbps * $0.000008)
moss.sound.effect
max($0.01, durationSeconds * $0.004 + maxNewTokens * $0.000008)
Balance request
curl -X GET "$API_BASE/api/v1/balance" \
-H "Authorization: Bearer $API_KEY"
Image generation request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-H "X-Idempotency-Key: job-001" \
-d '{
"task": "image.generate",
"input": {
"prompt": "cinematic fox astronaut",
"width": 1536,
"height": 1024,
"referenceImages": [
"https://cdn.example.com/input/style-ref.webp",
{
"dataBase64": "<base64-or-data-uri>",
"mimeType": "image/png"
}
]
}
}'Text-aware image design request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-H "X-Idempotency-Key: image-design-001" \
-d '{
"task": "image.design",
"input": {
"prompt": "A product launch poster with the headline ACME ROCKETS",
"preset": "quality",
"outputFormat": "webp",
"width": 1024,
"height": 1024
}
}'Text-aware image design request (JSON prompt)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-H "X-Idempotency-Key: image-design-json-001" \
-d '{
"task": "image.design",
"input": {
"prompt": {
"high_level_description": "A clean product launch poster for ACME ROCKETS.",
"style_description": {
"aesthetics": "minimal, premium, high-contrast",
"lighting": "even studio lighting",
"medium": "graphic_design",
"art_style": "modern editorial poster with crisp sans-serif typography",
"color_palette": ["#FFFFFF", "#111827", "#2563EB", "#F59E0B"]
},
"compositional_deconstruction": {
"background": "A clean white poster background with subtle depth.",
"elements": [
{
"type": "text",
"bbox": [90, 120, 250, 880],
"text": "ACME ROCKETS",
"desc": "Large, perfectly legible headline."
},
{
"type": "obj",
"bbox": [320, 200, 880, 800],
"desc": "A polished rocket product hero render."
}
]
}
},
"preset": "quality",
"width": 1024,
"height": 1024
}
}'Image edit request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "image.edit",
"input": {
"prompt": "turn this scene into a dramatic noir poster",
"sourceImage": "https://cdn.example.com/input/source.jpg",
"referenceImages": [
"https://cdn.example.com/input/style-ref.webp",
{
"dataBase64": "<base64-or-data-uri>",
"mimeType": "image/webp"
}
]
}
}'Video generation request (image-to-video)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "video.generate",
"input": {
"prompt": "slow cinematic push-in with floating particles",
"startFrameImage": {
"dataBase64": "<base64-or-data-uri>",
"mimeType": "image/png"
},
"finalFrameImage": "https://cdn.example.com/input/final-frame.webp",
"durationSeconds": 6,
"width": 1920,
"height": 1088,
"quality": "balanced",
"interpolationFps": 60,
"watermark": {
"image": "https://cdn.example.com/branding/logo.png",
"x": 1680,
"y": 968,
"width": 160,
"height": 80
}
}
}'Video generation request (text-to-video)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "video.generate",
"input": {
"prompt": "single continuous cinematic flythrough over a neon city at dawn",
"durationSeconds": 6,
"width": 1920,
"height": 1088,
"quality": "balanced",
"interpolationFps": 0
}
}'Video generation request (quality MiniMax H3)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "video.generate",
"input": {
"prompt": "Single continuous hero shot with restrained camera motion and strong identity retention.",
"startFrameImageUrl": "https://cdn.example.com/input/start-frame.webp",
"durationSeconds": 6,
"width": 2048,
"height": 1152,
"quality": "quality",
"interpolationFps": 60
}
}'Video generation request (multimodal H3 references)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "video.generate.reference",
"input": {
"prompt": "<Picture 1> defines the lead character. <Video 1> defines the handheld movement. <Audio 1> defines the pacing and sound palette.",
"referenceImageUrls": [
"https://cdn.example.com/input/character.webp"
],
"referenceVideoUrls": [
"https://cdn.example.com/input/camera-motion.mp4"
],
"referenceAudios": [
{
"url": "https://cdn.example.com/input/rhythm.opus"
}
],
"quality": "balanced",
"durationSeconds": 5,
"aspectRatio": "16:9",
"interpolationFps": 60
}
}'Video reference request (image + audio, H3 Turbo)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "video.generate.reference",
"input": {
"prompt": "Use <Picture 1> for identity and <Audio 1> to guide movement, timing, and generated sound.",
"referenceImageUrls": ["https://cdn.example.com/input/start-frame.webp"],
"referenceAudioUrls": ["https://cdn.example.com/input/reference-track.opus"],
"durationSeconds": 6,
"width": 1920,
"height": 1088,
"quality": "balanced",
"interpolationFps": 60
}
}'Video reference request (quality base H3)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "video.generate.reference",
"input": {
"prompt": "Use <Picture 1>, <Picture 2>, and <Picture 3> as high-fidelity visual references.",
"referenceImageUrls": [
"https://cdn.example.com/input/clip-a-last-03.webp",
"https://cdn.example.com/input/clip-a-last-02.webp",
"https://cdn.example.com/input/clip-a-last-01.webp"
],
"durationFrames": 243,
"width": 1280,
"height": 1280,
"interpolationFps": 60,
"quality": "quality",
"seed": 424242
}
}'Video combine request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "video.combine",
"input": {
"videoUrls": [
"https://cdn.example.com/output/clip-a.mp4",
"https://cdn.example.com/output/clip-b.mp4"
],
"overlapFrames": 3,
"frameRate": 60
}
}'Video resize request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "video.resize",
"input": {
"sourceVideoUrl": "https://cdn.example.com/output/clip-a.mp4",
"width": 1024,
"height": 1024,
"frameRate": 60,
"sourceDurationSeconds": 8
}
}'TTS request (Breeze-TTS-2)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "audio.speak",
"input": {
"provider": "breeze",
"mode": "voice_design",
"text": "Welcome aboard. Your journey begins now.",
"instructions": "A warm voice speaking clearly and naturally.",
"seed": 42
},
"idempotencyKey": "breeze-design-001"
}'TTS request (Gemini)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "audio.speak",
"input": {
"provider": "gemini",
"text": "We launch at dawn, hold formation, and bring everyone home safely.",
"voiceName": "Kore",
"instructions": "Speak with calm operational confidence and precise pacing."
}
}'TTS request (OmniVoice voice design)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "audio.speak",
"input": {
"text": "We launch at dawn, hold formation, and bring everyone home safely.",
"voiceDescription": "female, young adult, high pitch, american accent",
"language": "English",
"quality": "128k",
"seed": 123456
}
}'TTS voice clone request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "audio.speak",
"input": {
"mode": "voice_clone",
"text": "Keep the same character voice, but deliver this line with calm authority and clean pacing.",
"voiceDescription": "female, young adult, moderate pitch",
"referenceAudioUrl": "https://cdn.example.com/input/rin-voice-sample.opus",
"referenceTranscript": "Rin keeps her voice low. She measures every word before it lands.",
"language": "English",
"quality": "128k",
"seed": 424242
}
}'Audio annotation request (transcriptless ASR)
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "audio.annotation.reference",
"input": {
"sourceAudioUrl": "https://cdn.example.com/input/dialogue-take.opus",
"language": "en"
}
}'Music generation request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "ace.step.create",
"input": {
"tags": "anthemic ace step electro-pop, wide stereo synths, tight sidechained bass, punchy drum transients, strong vocal hooks, modern polished mix",
"lyrics": "[Intro]\nNeon rain on the avenue\n\n[Verse]\nWe were shadows in a crowded room\nNow the skyline sings our names\n\n[Pre-Chorus]\nHands up, hearts up, hold the line\n\n[Chorus]\nWe run through the midnight light\nTurn the static into fire tonight\n\n[Bridge]\nStrip it down, then build it higher\n\n[Final Chorus]\nWe run through the midnight light",
"qualityPreset": "high_quality",
"durationSeconds": 108,
"quality": "192k",
"language": "en"
}
}'Sound effect request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "moss.sound.effect",
"input": {
"prompt": "A sharp pistol shot in a dry canyon with a quick mechanical click and short tail echo.",
"durationSeconds": 2,
"quality": "128k",
"topK": 50,
"maxNewTokens": 1024
}
}'3D generation request
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"task": "trellis.generate",
"input": {
"sourceImageUrl": "https://cdn.example.com/input/boat_ref.png",
"qualityPreset": "balanced",
"targetFaceCount": 20000,
"textureSize": 1024,
"textureResolution": 512,
"maxViews": 4,
"sparseStructureSteps": 10,
"shapeSteps": 10,
"textureSteps": 10,
"meshResolution": 1024,
"remeshFillHoles": true,
"remeshFillHolesMaxPerimeter": 0.05,
"meshClusterConeHalfAngleRad": 55
}
}'Media job status request (summary)
curl -X GET "$API_BASE/api/v1/media/jobs?jobId=abc123" \
-H "Authorization: Bearer $API_KEY"
Media job status request (includeDetails=1)
curl -X GET "$API_BASE/api/v1/media/jobs?jobId=abc123&includeDetails=1" \
-H "Authorization: Bearer $API_KEY"
Media job request with completion callback
curl -X POST "$API_BASE/api/v1/media/jobs" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-H "X-Idempotency-Key: image-callback-001" \
-d '{
"task": "image.generate",
"completionCallbackUrl": "https://example.com/webhooks/media-job-terminal?token=replace-me",
"input": {
"prompt": "cinematic fox astronaut",
"width": 1024,
"height": 1024
}
}'Terminal callback payload example
{
"ok": true,
"task": "image.generate",
"externalJobId": "job_abc123",
"status": "completed",
"effectiveInput": {
"prompt": "cinematic fox astronaut",
"width": 1024,
"height": 1024
},
"inputAdjustments": [],
"billing": {
"usdCharged": "0.04"
},
"downloadableOutputUrls": [
"https://furgen-models.b-cdn.net/users/123/jobs/job_abc123/output_0001.webp"
],
"detailsIncluded": true,
"callbackType": "media_job_terminal"
}