Skip to main content
POST
Generate sound effects from text
The SFX endpoint turns a text prompt into a sound effect. It runs Stable Audio Open 1.0 (Stability AI) on ModelsLab GPUs and returns 44.1 kHz stereo audio.
For ElevenLabs Sound Effects on the same API key, use POST https://modelslab.com/api/v7/voice/sound-generation with model_id: eleven_sound_effect ($0.06 per generation). Price comparisons with other sound effects APIs are on modelslab.com/sound-effects-api-pricing.

Request

Make a POST request to the endpoint below with the required parameters.

Body

json

Response

When the GPU queue is clear, the request returns the finished file:
json
Otherwise the response has status: "processing", an eta in seconds, a fetch_result URL and the future_links where the file will appear. Keep the fetch_result URL (fetch replies do not repeat it) and POST to it with your key (see Fetch Voice) until status is success, failed or error, or pass a webhook to receive the result.

Body

application/json
key
string
required

API key required to authorize the request

prompt
string
required

Descriptive input specifying the type of sound effect to generate

duration
integer
default:8

Length of the sound effect in whole seconds

Required range: 3 <= x <= 15
output_format
enum<string>
default:mp3

Audio file format

Available options:
mp3,
wav,
flac
bitrate
enum<string>
default:320k

Audio bitrate

Available options:
128k,
192k,
320k
temp
boolean
default:false

Use temporary links for regions blocking storage access

webhook
string<uri>

URL to receive POST notification upon completion

track_id
integer

ID for webhook identification

Response

Sound effects generation response

status
enum<string>

Status of the voice generation

Available options:
success,
processing,
error
generationTime
number

Time taken to generate the audio in seconds

id
integer

Unique identifier for the voice generation

output
string<uri>[]

Array of generated audio URLs

Array of proxy audio URLs

Array of future audio URLs for queued requests

Array of audio URLs (voice cover response)

meta
object

Metadata about the audio generation including all parameters used

eta
integer

Estimated time for completion in seconds (processing status)

message
string

Status message or additional information

tip
string

Additional information or tips for the user

fetch_result
string<uri>

URL to fetch the result when processing

audio_time
number

Duration of the generated audio in seconds