Audio & video extraction

Gyrence can turn spoken content into either a timestamped transcript or structured JSON. The right resource depends on the output you need.

Choose the right resource

  • Need the spoken words with timestamps and speakers? Use Fetch with transcript: true. It returns kind: "transcript" with segments[].
  • Need facts, summaries, quotes, chapters, or other fields? Use Extract. It returns extractedJson shaped by your prompt and optional schema.
  • Need to upload a local audio or video file without writing code? Use Console → Extract for structured JSON from the file's transcript.

Supported inputs

  • YouTube: watch, Shorts, live, embed, youtu.be, and YouTube Music URLs when a caption track is available.
  • Podcast RSS: the newest feed item. Gyrence first reads a Podcasting 2.0 <podcast:transcript> file in WebVTT, SRT, or JSON format. If none is usable, it can transcribe the item's media enclosure.
  • Direct media URLs: MP3, M4A, AAC, WAV, OGG/OGA, Opus, FLAC, MP4, M4V, MOV, WebM, and MKV URLs, or URLs whose response declares an audio/video content type.
  • Local files: audio/video files up to 95 MB through Console → Extract.
Source priority

Existing source text always wins. Gyrence uses YouTube captions or the publisher's podcast transcript without a speech-recognition call. Speech recognition is used only for a direct media URL, a podcast enclosure without a usable transcript, or an uploaded file.

No-code: extract from a URL or file

  1. Open Console → Extract.
  2. Paste a YouTube, podcast-feed, or direct media URL. Alternatively, choose an audio/video file from your device.
  3. In what to extract, describe the result you need. Examples:
    • Summarize the main argument and return five supporting quotes with timestamps.
    • Return every company, product, person, and date mentioned.
    • Create chapters with a title, start time, and two-sentence summary.
  4. Optionally provide a JSON shape in schema.
  5. Select Run extract. File transcription may take a few minutes.

For URL inputs, Extract detects supported media automatically; there is no transcript switch to turn on. A natural-language description can also be resolved to a URL first, which adds Resolve credits.

HTTP: get a timestamped transcript

Set transcript: true on Fetch:

curl -X POST https://www.gyrence.com/api/v1/fetch \
  -H "Authorization: Bearer $GYRENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://www.youtube.com/shorts/VIDEO_ID",
    "transcript": true
  }'

When captions or a publisher transcript are available, the response completes immediately:

{
  "ok": true,
  "data": {
    "kind": "transcript",
    "status": "completed",
    "url": "https://www.youtube.com/shorts/VIDEO_ID",
    "source": "captions",
    "language": "en",
    "segments": [
      {
        "startMs": 0,
        "endMs": 2840,
        "speaker": null,
        "text": "Opening line of the video."
      }
    ],
    "dataClass": "public"
  }
}

source is captions, podcast_rss, or asr. Speaker labels are included when the source supplies them or speech recognition can separate speakers.

Poll speech-recognition jobs

Speech recognition can outlast a single request. If it is still running after the initial wait, Fetch returns:

{
  "ok": true,
  "data": {
    "kind": "transcript",
    "status": "processing",
    "url": "https://example.com/episode.mp3",
    "source": "asr",
    "jobId": "WORKSPACE_SCOPED_JOB_ID",
    "pollAfterMs": 3000,
    "dataClass": "public"
  }
}

Wait at least pollAfterMs, then send the jobId to the status endpoint using the same workspace API key:

curl -X POST https://www.gyrence.com/api/v1/fetch/transcript \
  -H "Authorization: Bearer $GYRENCE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"jobId":"WORKSPACE_SCOPED_JOB_ID"}'

The status endpoint returns another processing result or the completed transcript. Polling is free. A job ID is signed to its workspace and cannot be read with another workspace's key. A delivered speech-recognition transcript cannot be collected again.

MCP

The same workflows are available through the existing tools:

  • Call gyrence_fetch with { "url": "…", "transcript": true } for transcript segments.
  • Call gyrence_extract with a media URL, a prompt, and optional serialized JSON schema for structured output.
  • Local file upload is available in the Gyrence console, not through these URL-based MCP tools.

MCP uses the same handlers, response shapes, credits, and activity log as HTTP.

Credits

  • Fetch using existing YouTube captions or a podcast transcript: 1 credit.
  • Fetch using speech recognition: 10 credits when the job is submitted.
  • Transcript status polls: 0 credits.
  • Extract: the normal Extract price, 5 credits, or 7 only when browser rendering is used for a non-media page. Resolve-from-description adds 4 credits.

Limits and failure behavior

  • A YouTube video with no caption track returns not_found; Gyrence does not download or separate YouTube audio for speech recognition.
  • Podcast RSS processing targets the newest <item> in the feed.
  • Extract sends at most the first 12,000 characters of the transcript to the model. For long recordings, the structured result may cover only the beginning.
  • URL-based speech recognition submits the resolved public media URL directly. Gyrence does not download, extract, or transcode its audio first.
  • Uploaded files must identify as audio or video and must not exceed 95 MB.
  • Long speech-recognition work should use Fetch's asynchronous polling flow. The console's one-step Extract flow can time out on long files; retry with a shorter file or use transcript polling through the API/MCP.
  • Failed transcription returns an error or not_found; Gyrence never fabricates a transcript.

Data handling

Transcript inputs are treated as public data. Gyrence does not persist uploaded media or transcript text. Speech-recognition work is handled by AssemblyAI, and Gyrence requests deletion of the provider transcript after the completed result is delivered. See Subprocessors.