Azure AI Transcription SDK for Python
Client library for Azure AI Transcription (speech-to-text) with real-time and batch transcription.
Installation
pip install azure-ai-transcriptionEnvironment Variables
TRANSCRIPTION_ENDPOINT=https://<resource>.cognitiveservices.azure.com
TRANSCRIPTION_KEY=<your-key> # For key auth; not needed when using DefaultAzureCredential/TokenCredentialAuthentication & Lifecycle
๐ Two rules apply to every code sample below: 1. Two auth modes are supported:AzureKeyCredential(os.environ["TRANSCRIPTION_KEY"])for key-based auth, orDefaultAzureCredential()/ anyTokenCredentialfor Entra ID. PreferDefaultAzureCredentialin production; never hardcode credentials in code. 2. Wrap every client in a context manager so HTTP transports and sockets are released deterministically: - Sync:with <Client>(...) as client:- Async:async with <Client>(...) as client:Snippets may abbreviate this setup, but production code should always follow both rules.
Use subscription key authentication:
import os
from azure.core.credentials import AzureKeyCredential
from azure.ai.transcription import TranscriptionClient
with TranscriptionClient(
endpoint=os.environ["TRANSCRIPTION_ENDPOINT"],
credential=AzureKeyCredential(os.environ["TRANSCRIPTION_KEY"]),
) as client:
transcriptions = list(client.list_transcriptions())Transcription (Batch)
import os
from azure.core.credentials import AzureKeyCredential
from azure.ai.transcription import TranscriptionClient
with TranscriptionClient(
endpoint=os.environ["TRANSCRIPTION_ENDPOINT"],
credential=AzureKeyCredential(os.environ["TRANSCRIPTION_KEY"]),
) as client:
job = client.begin_transcription(
name="meeting-transcription",
locale="en-US",
content_urls=["https://<storage>/audio.wav"],
diarization_enabled=True,
)
result = job.result()
print(result.status)Transcription (Real-time)
import os
from azure.core.credentials import AzureKeyCredential
from azure.ai.transcription import TranscriptionClient
with TranscriptionClient(
endpoint=os.environ["TRANSCRIPTION_ENDPOINT"],
credential=AzureKeyCredential(os.environ["TRANSCRIPTION_KEY"]),
) as client:
stream = client.begin_stream_transcription(locale="en-US")
stream.send_audio_file("audio.wav")
for event in stream:
print(event.text)Best Practices
- Pick sync OR async and stay consistent. Do not mix
azure.xxxsync clients withazure.xxx.aioasync clients in the same call path. Choose one mode per module. - Always use context managers for clients and async credentials. Wrap every client in
with Client(...) as client:(sync) orasync with Client(...) as client:(async). For asyncDefaultAzureCredentialfromazure.identity.aio, also useasync with credential:so tokens and transports are cleaned up. - Enable diarization when multiple speakers are present
- Use batch transcription for long files stored in blob storage
- Capture timestamps for subtitle generation
- Specify language to improve recognition accuracy
- Handle streaming backpressure for real-time transcription
- Close transcription sessions when complete
Reference Files
| File | Contents |
|---|---|
| references/capabilities.md | Additional non-hero capabilities, operation-group coverage, and production checklists. |
| references/non-hero-scenarios.md | Dedicated non-hero examples for secondary/advanced scenarios. |
