Announced 6 Oct 2026 · Sources checked
What did OpenAI add to ChatGPT?
ChatGPT can now accept an audio recording as an uploaded file and work with its spoken content. OpenAI’s October 6 release note lists three core uses: producing a transcript, summarizing the recording and answering questions about it. That turns a meeting, lecture, interview or personal voice note into material that can be searched and reorganized inside the same conversation.
The feature is distinct from live voice conversations and from the separate developer API. It is a file-analysis workflow inside ChatGPT. For background on the speech-recognition technology that made modern transcription practical, see our guide to Whisper and speech models.
Who can use it, and which files work?
OpenAI says audio uploads are available on paid ChatGPT subscriptions and workspaces, including Enterprise, while the free plan does not include the feature. Availability can still vary with workspace settings, region, client and selected model, so an eligible account may not see exactly the same controls everywhere on launch day.
The supported list includes WAV, MP3 or MPEG audio, OGG or OGA, PCM, FLAC, AAC, M4A, audio-only WebM and audio-only MP4. OpenAI says files that the system identifies as video are not accepted through this audio-upload path, even if the container is WebM or MP4. Each file is limited to 512 MB.
How does the workflow handle long recordings?
The simplest flow is to attach an audio file and state the output you want: a verbatim-style transcript, a concise summary, an outline with decisions, or answers tied to the recording. A precise request helps ChatGPT choose the useful level of detail, but it does not change the underlying uncertainty in speech recognition.
When Data Analysis is available, OpenAI says ChatGPT may divide a long recording into smaller segments. That can make a large file more manageable, but it can also separate a sentence or topic from surrounding context. Upload and processing time depend on file size, duration, format and service load, and OpenAI notes that some requests can time out even when the file is below the formal size limit.
Treat the output as a draft that needs checking, especially when it will be quoted or used for decisions. Our practical AI-answer verification checklist applies here: compare important claims with the recording, not just with a fluent summary.
What can go wrong with transcription and speaker labels?
OpenAI explicitly warns that transcripts may contain errors and that performance varies by language. Proper names, technical vocabulary, accents, overlapping voices, background music and poor microphone placement are common sources of mistakes in any automatic speech-recognition system. A confident sentence can therefore contain a wrong name, date, quantity or negation.
Speaker identification is a separate limitation. The help page says ChatGPT may not reliably distinguish or label speakers. A meeting summary that attributes a commitment to the wrong person can be more damaging than a misspelled word, so teams should verify speaker turns before copying action items into a system of record.
These failures resemble other forms of AI hallucination and overconfident output: readability is not proof of accuracy. For legal, medical, financial or disciplinary use, keep the original recording and use a review process appropriate to the stakes.
What should users know about privacy and retention?
An audio file can contain voices, health details, workplace discussions, customer information and other sensitive material. Before uploading, users should confirm that they have permission to record and process the conversation, follow local law and workplace policy, and remove material that the service does not need. The new upload control does not itself establish consent.
OpenAI’s help page directs users to their plan’s data controls and retention rules. Consumer and business offerings can differ, and Enterprise administrators may impose additional workspace settings. The safe question is not merely whether ChatGPT can read the file, but whether the recording is appropriate for the selected account and whether the resulting transcript should be stored or shared.
Our overview of AI API data retention explains why storage duration, training settings, deletion behavior and administrator controls should be checked separately. The ChatGPT product workflow and API audio endpoints also have different formats, limits and policies.
Where is the feature most useful—and where is it not?
The strongest uses are low-friction first drafts: converting a clear interview into searchable notes, extracting themes from a lecture, locating a topic in a long voice memo or building a meeting recap that a participant then reviews. Asking for timestamps, uncertain passages and a list of names to verify can make that review faster when the product returns enough detail.
The feature should not be treated as an automatic official record. OpenAI has not claimed perfect word accuracy, reliable diarization or consistent performance across every language and recording condition. Teams comparing it with a specialist transcription service should test representative audio, measure corrections per minute and inspect how each tool handles speakers, jargon and sensitive data.
In practice, audio uploads remove a conversion step and make recordings easier to explore, but the useful output is a reviewed transcript or summary—not the first generated text. Keep the source audio, mark uncertain passages and have a person approve anything that will be published, cited or assigned.
Common questions
Can free ChatGPT accounts upload audio files?
No. OpenAI’s current help page says audio uploads are available to paid subscriptions and workspaces; the free plan does not include them.
Does ChatGPT reliably identify each speaker?
No. OpenAI says speaker identification may be unreliable. Verify every attribution before using a transcript for decisions, quotations or action items.
Is an audio upload the same as using OpenAI’s audio API?
No. This is a ChatGPT product feature. The API has separate endpoints, supported formats, request limits, pricing and data-handling rules.
What to remember
OpenAI’s new audio-upload workflow makes transcription, summaries and follow-up questions available inside ChatGPT for paid users and eligible workspaces. It supports common formats and files up to 512 MB, while OpenAI cautions that processing can fail and transcripts, languages and speaker labels can be inaccurate. Use it for reviewed drafts, not as an unquestioned record.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





