Turn recordings into original-language text and usable subtitles
whisper-1 is OpenAI's speech-to-text model, corresponding to Whisper large-v2 when the official API was released. It is suitable for organizing speech from interviews, courses, voice messages, and audiovisual materials into original-language text, with text, JSON, or subtitle formats selected according to the workflow. On this platform, you can submit files through the audio transcription entry point to integrate recording processing into content production and document organization workflows.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Clarify capacity, input/output, and invocation methods before selecting a model.
Native model relationship
Whisper large-v2; API invocation ID is whisper-1
Input and output
Audio file input, original-language text transcription output
Native file formats
m4a、mp3、mp4、mpeg、mpga、wav、webm
Invocation endpoint
POST /v1/audio/transcriptions; submit binary file
Output format options
json、text、srt、verbose_json、vtt; default is json
Transcription controls
Can pass language、prompt; temperature defaults to 0
Native format details and invocation parameters are listed separately; this entry point focuses on file transcription and is not equivalent to real-time voice conversation.
Core capabilities
Learn what whisper-1 can bring to your work.
A transcription path that preserves the original language
whisper-1's transcription task converts spoken audio into original-language text rather than translating it into English by default. This is suitable for interviews, courses, and dictated materials that need to retain the original wording. The generated text can serve as a basis for subsequent search, excerpting, and editing, while summarization and rewriting should be handled as separate processing steps.
Deliver text and subtitles separately
The same transcription workflow can select output formats according to the use case: text is suitable for direct entry into an editor, JSON makes text easy for programs to read, and SRT and VTT are suitable for subtitle production. Format selection should be designed around the final deliverable, avoiding saving plain text first and then additionally rebuilding subtitle structures and processing workflows.
Organize calls around files
Calls are centered on audio files, so recordings do not need to be converted into chat messages. Requests can include language and prompt text to help applications organize transcription context; prompts should be prepared around the recording content rather than used as free-form creative instructions. This makes it easier to connect uploading, transcription, storage, and human proofreading into a stable workflow.
Applicable Scenarios
Start with specific tasks to find where the model can be useful.
Interview and Oral Material Organization
Submit interview recordings or oral materials as files, choose text or JSON output, and obtain transcripts that are easy to search and edit. Editors can use them to locate topics, organize quotations, and listen back to verify names, organization names, and key original statements; article structure and viewpoint extraction can be handled separately after transcription is complete.
Course and Video Subtitle Production
Choose SRT or VTT output for course recordings or video materials, and send the results into subtitle editing or playback workflows. Before delivery, check sentence breaks and the correspondence between text and timing against the original material; the model handles speech transcription, while subtitle layout, visual styling, and final video export are still completed by production tools.
Voice Message Archiving
Convert user voice messages into text for storage in ticketing systems or knowledge bases. Applications can read the text in JSON and associate the transcription with the original audio, making it easier for staff to review and listen back. Classification, priority assessment, and reply generation should be handled in separate processing steps rather than treating transcription results directly as business decisions.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Choose It When You Need a Text Transcript
If the core requirement is to process existing recordings and deliver original-language text or subtitles, whisper-1 has clear task boundaries. It is not a general question-answering model, nor is it a speech synthesis model: it will not read text aloud as audio, and it should not be expected to handle all the work of automatically writing meeting notes. Complete transcription first, then follow with summarization or editing workflows to make it easier to verify the original wording.
Distinguish by Version and Task
whisper-1 is the API invocation name, while large-v2 is the corresponding native version at the time of official release; it should not be understood as a general term for all newer Whisper versions. When comparing other transcription models, use the same recordings to check proper nouns, sentence breaks, and required output formats; if the goal is English translation, also distinguish between transcription in the original language and independent translation tasks.
Get Started
From a small-scale task to formal integration.
01
Prepare the Task and Materials
Define the objective, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Testing Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Boundaries
Before formal use, understand the output quality and scope of capabilities.
Transcribed text is not an automatically generated summary, meeting conclusion, or record of speaker identity. When multiple-speaker recordings require distinguishing speakers, arrange a separate speaker-labeling step; do not infer the attribution of each sentence solely from the transcript, and preserve the correspondence with the original audio for important quotations.
Input should use audio files; do not submit PDFs, chat text, or web links directly as binary files. When calling through MCP's audio URL tool, distinguish this from the HTTP file upload process as well; file preparation and reading methods depend on the client used.
File transcription is not equivalent to a real-time speech service that recognizes speech while recording, and subtitle output does not mean that all editing work is complete. When real-time interaction, precise timing alignment, or formal subtitle delivery is needed, test the workflow against the actual audio and schedule review listening and subtitle editing.
Frequently Asked Questions
Answers to common questions about using whisper-1.
What is the relationship between whisper-1 and Whisper large-v2?
whisper-1 is the model name used when calling the API, and OpenAI's Whisper API release notes map it to large-v2. They describe the invocation name and native version respectively; whisper-1 should not be treated as a generic alias for all Whisper versions, nor should this be taken to mean that it is the latest version.
Will Chinese recordings automatically become English?
The task for the audio transcription endpoint is to output text in the original language, not to translate it into English by default. Officially, transcription in the original language and translation into English are treated as separate tasks. Therefore, Chinese recordings should be processed using the Chinese transcription workflow; if an English transcript is needed, a separate translation step should be arranged.
Can it generate SRT or VTT subtitles directly?
The call parameters provide srt and vtt output options for subtitle workflows; for plain-text tasks, text or json can be selected. After subtitles are generated, the content, line breaks, and timing should still be checked against the audio before handing them off to subtitle editing or playback tools; they should not be regarded as a finished product directly.
Can prompt make it rewrite a recording into an article?
prompt is hint text in a transcription request and should not be treated as a general-purpose writing instruction. The core deliverable of whisper-1 remains an audio transcript. When an article, summary, or action items are needed, it is recommended to retain the original transcription result and then pass it to a separate text-processing step, avoiding confusion between the original words and processed content.
Should I upload a file or submit an audio URL?
When calling /v1/audio/transcriptions directly, you need to submit a binary file and use whisper-1. If using an MCP client, you can choose its tool for transcribing audio from a URL. The inputs are organized differently for the two methods; do not submit a URL string directly as the binary file field in HTTP.
Model information · Updated: 2026-10-01. See the API and pricing sections for call parameters and billing rules.