Tools provides a queue-based Whisper transcription service for media URLs and uploaded audio/video files.
This guide focuses on the public contract: what the feature does, how users interact with it, which endpoints exist, and what clients should expect in requests and responses.
Whisper in Tools can:
The ordinary Whisper UI and most authenticated API endpoints use:
POST /api/account/loginThe re-transcription mutation has an explicit split: browser users submit POST /whisper/jobs/{jobId}/retranscription with the signed-in web session and normal CSRF protection, while machine clients submit POST /api/whisper/jobs/{jobId}/retranscribe with JWT bearer auth. The machine mutation does not use the browser session as a fallback identity.
User permission requirements:
whisper.use for ordinary queue accesswhisper.manage for full-queue/admin actions such as run-now and all-user visibilityprovider_openai when transcript analysis/translations should run for a non-admin userTools now also exposes a separate server-to-server transcription API for token-based integrations.
Auth requirements:
whisper.api (the built-in generator still creates a provider_whisper_api token row for convenience)Authorization: Bearer YOUR_API_TOKENwhisper.apiwhisper.use)Legacy X-Api-Key or apikey transport may still exist for backwards compatibility, but new integrations should use the Authorization header.
Whisper jobs are processed asynchronously.
Job statuses:
queueddownloadingtranscribingfinalizingcompletedfailedJobs can also expose a queue origin:
queue_channel="web"queue_channel="api"Signed-in queue/detail views and authenticated /api/whisper/jobs* payloads can therefore show whether a job came from the regular user queue or the token-authenticated API queue.
Admin-owned jobs are prioritized ahead of non-admin jobs when queued work is claimed.
Administrators can also restrict which models the main Tools host is allowed to execute locally from the Whisper administration runtime settings. The local allowlist defaults safely to small until an administrator changes it. Jobs that request another model remain eligible for compatible authenticated remote workers; the Tools host skips them instead of silently running a blocked model locally. Remote workers keep their own advertised model capabilities and are not reduced by the local-host allowlist.
On an existing job, administrators can explicitly select Automatic selection, Tools local, or a configured named remote worker as the execution target. Automatic selection remains fully capability-gated. Exact targets never fall back to another runner. A named remote worker is also an explicit force attempt: on that worker's next valid current-contract claim poll, Tools can assign the exact queued job across ordinary advertised model or speaker-diarization mismatches. For a URL-origin job, a worker that advertises URL support can still receive the source URL directly. If the selected remote path cannot consume external URLs, Tools first downloads the source itself, keeps the verified media on the Whisper host, returns the job to the queue, and gives the worker a lease-bound Tools-hosted media download instead. The worker then executes the original transcription requirements or reports a normal retryable/terminal failure instead of leaving the job silently queued. Force execution never bypasses worker authentication, current contract-version compatibility, lease ownership, retry limits or required source/media integrity, and Tools local keeps its local enabled/model/diarization safety settings. Offline remotes may still be selected in advance. The selector distinguishes forceable capability mismatches from protocol/media conditions that cannot be forced. (#1727, #1732)
/whisperThe signed-in queue UI lets users:
For local uploads, multiple files can be selected at once. Tools uploads them one at a time and creates a separate Whisper job for each file with the same selected model, language, analysis, translation, and diarization settings.
During transfer, the page shows both overall batch transfer progress and separate progress/status for each file. After a file transfer completes, its row shows that Tools is waiting for the queue job to be created. If one file fails, the remaining files continue and successful jobs stay in the queue. The effective upload limit applies per file.
The authenticated API remains a single-file contract: media_file means one file per request. Multi-file behavior in /whisper is client-side batching on top of the existing enqueue API.
The queue also shows a Whisper workers card for configured remote transcription workers. Worker availability and last-seen data refresh through authenticated AJAX every 15 seconds while the tab is visible, and Refresh worker status performs the same refresh without reloading the page. Automatic worker polling pauses while the tab is hidden and refreshes immediately when it becomes visible again. If a refresh fails temporarily, the card keeps the last valid worker snapshot and reports the refresh problem instead of incorrectly replacing the list with an empty/offline state.
For URL jobs, Tools normally honors the configured yt-dlp proxy. If that proxy itself cannot be resolved or reached, the runner retries the same download directly so a temporary proxy outage does not immediately fail the job. Installations that require every media request to use the proxy can enable strict proxy mode; in that mode no direct retry is made.
Facebook extraction can return an anonymous "cannot parse data" result even when a saved signed-in session is available. Tools now treats that response as a reason to continue through the configured cookie sources instead of stopping after the anonymous attempts.
/whisper/jobs/{jobId}The signed-in detail page shows:
Completed jobs can create a public transcript share page.
When a Whisper job reaches a final completed or failed state and the owner has an email address, the owner report links back to the public Tools job page and never uses a local-only job address.
The report also contains a speaker-diarization summary. It states whether diarization was requested or disabled, shows its current or final status, and includes the provider, detected speaker count, labelled segment count, and safe warning or error text when those values are available.
When the source media for a completed job is still retained, the transcript owner can choose Re-transcribe and select another model that is actually supported by the current Whisper runtime. The selected model must differ from the current pass. The existing completed pass is archived before the new pass enters the normal queue, so trying a heavier or otherwise different model does not destroy the previous transcript.
The job detail page keeps the ordinary job as the current/primary pass for backwards compatibility and shows earlier completed passes as read-only history. Each archived pass records the model and language used, its transcript, timestamped transcript segments, and the transcript-dependent analysis, translations, diarization and speaker-editor state that belonged to that pass. The new primary pass starts without those old transcript-dependent results; they can be generated again from the new transcript instead of being silently carried across models.
Re-transcription is owner-only even when another administrator can view the job. It is unavailable while the current pass is not completed, while translation or speaker-diarization work is still pending, or after the retained source media has been released. The UI shows the reason instead of offering a broken action. URL jobs reuse the retained local copy selected for the new pass; Tools does not silently download the external URL again. With the current remote-worker contract, retained-media URL re-transcriptions therefore stay on the local runner so a worker cannot accidentally re-fetch the original URL. Their replacement-model choices are limited to the current local Tools-host allowlist, and the action is unavailable when no different locally allowed model exists.
Model size is a capacity, compute and accuracy/robustness tradeoff rather than a simple vocabulary-size choice. Larger models generally require more resources and can perform better on difficult, multilingual or noisy speech. turbo is an optimized model derived from large-v3 and is useful when a stronger speed/accuracy balance is preferred.
For uploaded media, Tools tracks whether the current job title is still only the original upload filename. When the transcript completes, a user with OpenAI access can have that filename fallback replaced automatically with a short title based on the transcript content. A title explicitly supplied or edited by the user is never replaced, and a title-generation failure does not change an otherwise successful Whisper job into a failed job.
The completed transcript view also offers a dedicated title-suggestion action beside the transcript. It reuses the existing transcript metadata AI helper but fills only the editable title field. The suggestion is a preview: it is not persisted until the transcript owner reviews it and explicitly saves the metadata form. The existing title-and-description auto-fill action remains available separately.
Automatic and manually requested title-generation operations are written to the dedicated Whisper audit log as operational events. Transcript text and generated title content are not stored in that audit context.
While Whisper is transcribing, the live transcript view is updated through the same job-detail polling as the workflow log. Local runs and current remote workers publish completed segment evidence as it becomes available; remote progress carries bounded cumulative text plus timestamped segments, so Tools can derive progress and observed throughput from the latest completed audio position instead of treating a fixed worker stage marker as processed audio. Text appears one completed Whisper segment at a time, not word by word. The card keeps a bounded scroll area so a long transcript does not keep expanding the page layout. Incoming segments keep the view at the newest text only while the scrollbar is already at the bottom. Scrolling upward pauses that follow behavior until the scrollbar returns to the bottom.
The Live progress card now shows a prominent, color-coded runner state. It distinguishes a normally running job from a delayed heartbeat, a runner that has likely stopped, a stale job that is likely dead, a queued job waiting for a runner, and a finished job. The exact last heartbeat and last output times remain visible under the state.
Speaker diarization has a separate progress indicator in the speaker section. The transcription progress remains at 100% once the transcript itself is complete, while a queued or running diarization rerun is shown independently. The detail page follows the structured pyannote progress events already stored in the timestamped workflow log, so a measurable provider stage can show its current stage name, completed/total work, stage percentage, observed elapsed time, and age of the latest provider update. That percentage is explicitly stage-local and is never presented as the percentage for the whole diarization run. When the current provider stage has no trustworthy completed/total measurement, the diarization bar remains deliberately indeterminate instead of inventing a value. Queue, worker-claim, and provider-start markers are treated as boundaries for a new rerun, so retained provider events from the previous run are discarded immediately. If the transcript owner opens the page while a rerun is already pending or running, the rerun control remains available in a disabled state and confirms the current backend capability before becoming actionable again after a terminal transition. A terminal completed diarization still shows the whole-run indicator as 100%. If older stored jobs are missing timestamped JSON segments, Tools can also recover timings from generated subtitle artifacts before diarization without replacing the saved transcript, analysis, or translations.
If an active job has no runner heartbeat for 90 seconds, the page marks it as likely stopped. This warning is separate from the longer automatic stale-recovery threshold, so the UI can warn early without changing when a job is automatically reclaimed.
Workflow-log timestamps and all job-detail timestamps are shown explicitly in Europe/Stockholm, including Swedish standard time and daylight-saving time. The estimated finish uses the same time zone and includes an approximate remaining time. It prefers completed transcript segment timing when available, falls back to the current transcribing progress before that, and is recalculated by the existing polling as processing speed changes.
Each new run also mirrors the evolving text into a dedicated TXT snapshot after every completed segment. The job page shows the snapshot's storage and physical paths in the storage path trace. Long path values wrap inside their own rows so they do not overlap adjacent details.
The main Tools host keeps its managed Whisper virtualenv plus bundled Node/Deno yt-dlp runtimes under the configured Whisper storage root. Normal full gitpull and repository maintenance preserve those managed executable bits and revalidate the configured WHISPER_BIN, Python interpreter and JavaScript runtimes. php artisan whisper:doctor should therefore report executable/resolvable managed binaries after maintenance without requiring manual chmod; external runtime overrides remain operator-owned and are not modified automatically.
Completed transcripts can be exposed through a tokenized public page under:
/shared/whisper/transcript/{token}The share page is intended for reading/transcript sharing, not queue administration.
For token-authenticated API submissions, Tools can now create that share automatically when the transcript completes successfully, and the callback payload includes the direct share URL.
/api/whisper/*)Most endpoints in this group accept signed-in web-session or JWT authentication and do not use the dedicated Whisper API token. The exception is the machine re-transcription mutation POST /api/whisper/jobs/{jobId}/retranscribe, which requires JWT bearer authentication and deliberately does not use a browser session as its caller identity. Browser re-transcription uses the separate CSRF-protected POST /whisper/jobs/{jobId}/retranscription action.
GET /api/whisper/statusReturns queue counters and capability flags.
Typical response shape:
{
"ok": true,
"summary": {
"queued": 3,
"processing": 1,
"completed": 21,
"failed": 2
},
"can_manage_all": false,
"config": {
"enabled": true,
"default_model": "large",
"upload_max_mb": 64,
"upload_limit": {
"configured_mb": 200,
"php_upload_max_mb": 64,
"php_post_max_mb": 128,
"effective_max_mb": 64,
"effective_max_label": "64 MB",
"limited_by_php": true
},
"ytdlp_configured": true
}
}
upload_max_mb now reflects the practical/effective limit for uploaded Whisper media on the current host, and additive config.upload_limit can explain when PHP upload/body limits are lower than Whisper's own configured cap.
GET /api/whisper/jobs?limit=100Returns visible Whisper jobs for the authenticated user.
POST /api/whisper/jobsQueues a new Whisper job.
Supported request styles:
source_urlmultipart/form-data with media_fileImportant rule:
source_url or media_file, not bothmedia_file is one file per request; the web UI's multi-file mode sends separate requests and creates separate jobs422 validation error under media_file instead of only the generic "failed to upload" wordingExample JSON body:
{
"source_url": "https://example.com/audio.mp3",
"source_label": "Interview with customer",
"source_note": "Recorded support follow-up call.",
"model": "large",
"language": "sv",
"analysis_language": "sv",
"translation_target_languages": ["en"]
}
GET /api/whisper/jobs/{jobId}Returns one visible Whisper job. The job returned here remains the current/primary transcription pass even after re-transcription, so existing clients that expect one transcript continue to use the same contract.
Additive job fields now include:
queue_channelqueue_channel_labelsource_typesource_labelsource_notesource_mimesource_size_bytessource_duration_secondssource_duration_humanstage_labelstage_detailruntime_log[]livenessanalysistranslations[]diarizationsharecallback (primarily relevant for API-queue jobs)The additive liveness.state value can be inactive, active, quiet, unresponsive, stale, or suspect. unresponsive means the heartbeat has been absent long enough that the runner likely stopped, while stale means the existing automatic stale-recovery threshold has also been reached.
GET /api/whisper/jobs/{jobId}/revisionsReturns re-transcription state for one visible job: the explicit current/primary pass, supported model choices, whether the owner can start another pass, a human-readable unavailable reason when not, and earlier completed revision history. Historical revisions can include the transcript plus the analysis, translations, diarization and speaker-editor state that belonged to that archived pass.
Reading follows the ordinary Whisper job-visibility boundary. Starting a new pass remains owner-only.
POST /api/whisper/jobs/{jobId}/retranscribeQueues a new transcription pass from the retained source media.
Authentication:
POST /api/account/login as Authorization: Bearer YOUR_JWTExample body:
{
"model": "large"
}
Guardrails:
The current completed pass is archived transactionally before the job is queued again. Its transcript-dependent analysis, translations, diarization and speaker state remain attached to that history item, while the new primary pass starts clean. A new pass is also a new terminal-notification cycle for the owner.
POST /api/whisper/jobs/{jobId}/analyzeRuns transcript analysis for a completed transcript.
Guardrails:
POST /api/whisper/jobs/{jobId}/cancelRequests cooperative cancellation for an actively processing job.
POST /api/whisper/jobs/{jobId}/restartQueues a failed/queued job for retry.
DELETE /api/whisper/jobs/{jobId}Deletes a non-processing job.
POST /api/whisper/run-nowAdmin/manager helper endpoint.
Request body can include:
{
"limit": 1,
"reset_failed": true
}
/api/whisper/transcribe/*)This is the dedicated server-to-server callback API.
GET /api/whisper/transcribe/statusReturns queue counters for the token-authenticated API queue channel.
GET /api/whisper/transcribe/jobs?limit=100Returns visible jobs from the API queue channel.
GET /api/whisper/transcribe/jobs/{jobId}Returns one visible API-queue job.
POST /api/whisper/transcribeQueues a new token-authenticated Whisper job.
Required field:
callback_urlSupported submission styles:
source_urlmedia_fileUpload validation guidance:
422 with errors.media_file[] explaining the effective Whisper upload limitmedia_file validation path is also used for partial uploads, missing temp-folder failures, write failures, and other PHP upload transport errors before the job is queuedThe token API accepts the same additive metadata as the ordinary queue endpoint, including:
source_labelsource_notemodellanguageanalysis_languagetranslation_target_languages[]disable_diarizationSpeaker diarization is requested by default when the feature is available. Omit disable_diarization (or send a false value) to keep diarization enabled; send a truthy value only when that job should skip speaker diarization. The /whisper checkbox is therefore an opt-out and is unchecked by default. Guest Whisper jobs use the same default.
Example JSON body:
{
"source_url": "https://example.com/audio.mp3",
"callback_url": "https://api.example.test/whisper/callback",
"source_label": "Customer interview",
"source_note": "Transcribe and send the final result back to our integration.",
"model": "large",
"language": "en",
"analysis_language": "en",
"translation_target_languages": ["sv"]
}
Example success response:
{
"ok": true,
"message": "Whisper API job queued. A callback will be sent when the job reaches a terminal state.",
"job": {
"id": 123,
"queue_channel": "api",
"queue_channel_label": "API queue",
"status": "queued",
"callback": {
"url": "https://api.example.test/whisper/callback",
"status": "pending",
"http_status": null,
"last_attempt_at": null,
"delivered_at": null,
"error": null
},
"share": null
}
}
When a token-authenticated Whisper API job reaches terminal completed or failed, Tools sends one JSON POST to the submitted callback_url.
Callback envelope:
{
"ok": true,
"event": "whisper.job.completed",
"job": {
"job_id": 123,
"status": "completed",
"status_label": "Completed",
"queue_channel": "api",
"queue_channel_label": "API queue",
"source": "Customer interview",
"model": "large",
"language": "en",
"job_url": "https://tools.example.test/whisper/jobs/123",
"share_url": "https://tools.example.test/shared/whisper/transcript/example-token-redacted",
"transcript_text": "...",
"analysis_text": "...",
"translations": [],
"share": {
"url": "https://tools.example.test/shared/whisper/transcript/example-token-redacted"
}
}
}
Failure callbacks use event="whisper.job.failed" and can include failure_error instead of transcript/share data.
Client guidance:
job.job_idTypical error classes:
401 unauthenticated / token rejected403 missing permission404 job not found or not visible to the caller422 validation or business-rule failure429 throttled5xx temporary backend/provider failureWhisper API routes use a general throttle:120,1 policy.
Clients should still implement normal backoff for repeated polling or transient failures.
Authorization: Bearer YOUR_API_TOKENcallback_url as required for token-authenticated submissionsqueue_channel and queue_channel_label in operator/debug UIstranscript_text as the primary result and speaker_aware_transcript as additive helper outputshare.url as public access and handle it carefullyA completed transcript can queue speaker diarization afterwards even when diarization was disabled during the original transcription. Tools reuses the retained local audio/video file. When timestamped transcript segments are already stored, the rerun goes directly to diarization. If those older segment records are missing, the worker regenerates only the timestamped segments from the retained media, stores them on the existing job, and then runs diarization without replacing the saved transcript text, analysis, or translations.
Post-transcription diarization therefore requires the source media to still exist in Whisper storage, but saved transcript segments are no longer required up front. If the media has already been purged, the request is rejected clearly instead of creating a broken queue item. If segment recovery fails later, the completed transcript remains available and only the diarization attempt is marked failed. Enabling diarization afterwards also persists the preference on the job, so later runs behave as ordinary diarization reruns. During a pending or running rerun, the transcript owner's rerun control remains disabled; when the run reaches completed, failed, unavailable, or skipped, the UI rechecks the current backend capability before enabling that control. This allows a page opened mid-run to recover the action without a manual reload while keeping backend authorization and current job state authoritative.
Whisper job details separate worker progress from transcript evidence. The page distinguishes a worker that is still transcribing, transcript text that has actually reached Tools, a worker that has finished transcription and moved into post-transcription speaker diarization, and a final transcript that is actually stored. A high progress percentage is not treated as proof that Tools has the transcript. Remote completion stores the database transcript as the authoritative result; if only the auxiliary transcript TXT artifact cannot be written, the completed transcript remains available and the artifact problem is reported separately as degraded storage.
On a completed transcript with retained source media, Run / re-run diarization reuses the existing transcript and source media; it does not run Whisper again or replace the saved transcript. After a diarization result reaches completed, failed, unavailable, or skipped, another attempt may be queued. While diarization is genuinely pending or running, the action stays unavailable to prevent duplicate execution.