Audio & video extraction
Gyrence can turn spoken content into either a timestamped transcript or structured JSON. The right resource depends on the output you need.
Choose the right resource
- Need the spoken words with timestamps and speakers? Use Fetch with
transcript: true. It returnskind: "transcript"withsegments[]. - Need facts, summaries, quotes, chapters, or other fields? Use Extract. It returns
extractedJsonshaped by your prompt and optional schema. - Need to upload a local audio or video file without writing code? Use Console → Extract for structured JSON from the file's transcript.
Supported inputs
- YouTube: watch, Shorts, live, embed,
youtu.be, and YouTube Music URLs when a caption track is available. - Podcast RSS: the newest feed item. Gyrence first reads a Podcasting 2.0
<podcast:transcript>file in WebVTT, SRT, or JSON format. If none is usable, it can transcribe the item's media enclosure. - Direct media URLs: MP3, M4A, AAC, WAV, OGG/OGA, Opus, FLAC, MP4, M4V, MOV, WebM, and MKV URLs, or URLs whose response declares an audio/video content type.
- Local files: audio/video files up to 95 MB through Console → Extract.
Existing source text always wins. Gyrence uses YouTube captions or the publisher's podcast transcript without a speech-recognition call. Speech recognition is used only for a direct media URL, a podcast enclosure without a usable transcript, or an uploaded file.
No-code: extract from a URL or file
- Open Console → Extract.
- Paste a YouTube, podcast-feed, or direct media URL. Alternatively, choose an audio/video file from your device.
- In what to extract, describe the result you need. Examples:
Summarize the main argument and return five supporting quotes with timestamps.Return every company, product, person, and date mentioned.Create chapters with a title, start time, and two-sentence summary.
- Optionally provide a JSON shape in schema.
- Select Run extract. File transcription may take a few minutes.
For URL inputs, Extract detects supported media automatically; there is no transcript switch to turn on. A natural-language description can also be resolved to a URL first, which adds Resolve credits.
HTTP: get a timestamped transcript
Set transcript: true on Fetch:
curl -X POST https://www.gyrence.com/api/v1/fetch \
-H "Authorization: Bearer $GYRENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://www.youtube.com/shorts/VIDEO_ID",
"transcript": true
}'When captions or a publisher transcript are available, the response completes immediately:
{
"ok": true,
"data": {
"kind": "transcript",
"status": "completed",
"url": "https://www.youtube.com/shorts/VIDEO_ID",
"source": "captions",
"language": "en",
"segments": [
{
"startMs": 0,
"endMs": 2840,
"speaker": null,
"text": "Opening line of the video."
}
],
"dataClass": "public"
}
}source is captions, podcast_rss, or asr. Speaker labels are included when the source supplies them or speech recognition can separate speakers.
Poll speech-recognition jobs
Speech recognition can outlast a single request. If it is still running after the initial wait, Fetch returns:
{
"ok": true,
"data": {
"kind": "transcript",
"status": "processing",
"url": "https://example.com/episode.mp3",
"source": "asr",
"jobId": "WORKSPACE_SCOPED_JOB_ID",
"pollAfterMs": 3000,
"dataClass": "public"
}
}Wait at least pollAfterMs, then send the jobId to the status endpoint using the same workspace API key:
curl -X POST https://www.gyrence.com/api/v1/fetch/transcript \
-H "Authorization: Bearer $GYRENCE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"jobId":"WORKSPACE_SCOPED_JOB_ID"}'The status endpoint returns another processing result or the completed transcript. Polling is free. A job ID is signed to its workspace and cannot be read with another workspace's key. A delivered speech-recognition transcript cannot be collected again.
MCP
The same workflows are available through the existing tools:
- Call
gyrence_fetchwith{ "url": "…", "transcript": true }for transcript segments. - Call
gyrence_extractwith a media URL, aprompt, and optional serialized JSONschemafor structured output. - Local file upload is available in the Gyrence console, not through these URL-based MCP tools.
MCP uses the same handlers, response shapes, credits, and activity log as HTTP.
Credits
- Fetch using existing YouTube captions or a podcast transcript: 1 credit.
- Fetch using speech recognition: 10 credits when the job is submitted.
- Transcript status polls: 0 credits.
- Extract: the normal Extract price, 5 credits, or 7 only when browser rendering is used for a non-media page. Resolve-from-description adds 4 credits.
Limits and failure behavior
- A YouTube video with no caption track returns
not_found; Gyrence does not download or separate YouTube audio for speech recognition. - Podcast RSS processing targets the newest
<item>in the feed. - Extract sends at most the first 12,000 characters of the transcript to the model. For long recordings, the structured result may cover only the beginning.
- URL-based speech recognition submits the resolved public media URL directly. Gyrence does not download, extract, or transcode its audio first.
- Uploaded files must identify as audio or video and must not exceed 95 MB.
- Long speech-recognition work should use Fetch's asynchronous polling flow. The console's one-step Extract flow can time out on long files; retry with a shorter file or use transcript polling through the API/MCP.
- Failed transcription returns an error or
not_found; Gyrence never fabricates a transcript.
Data handling
Transcript inputs are treated as public data. Gyrence does not persist uploaded media or transcript text. Speech-recognition work is handled by AssemblyAI, and Gyrence requests deletion of the provider transcript after the completed result is delivered. See Subprocessors.
