语音识别接口接收音频文件,并使用设备上运行的语音识别模型返回文本结果。本文示例使用 AI Pyramid 的 SenseVoice 模型。
调用接口前,需要在设备上安装语音识别模型和示例所需的工具。模型支持情况请参考模型库。
apt install llm-model-sense-voice-small-10s-ax650 ffmpeg 工具。apt install ffmpeg systemctl restart llm-openai-api 以下示例在 PC 端使用 OpenAI Python SDK 提交音频文件。调用远程设备时,请将 base_url 中的地址替换为设备的实际 IP 地址。
from openai import OpenAI
client = OpenAI(
api_key="sk-",
base_url="http://192.168.20.186:8000/v1"
)
audio_file = open("speech.mp3", "rb")
transcript = client.audio.transcriptions.create(
model="sense-voice-small-10s-ax650",
file=audio_file
)
print(transcript) | 参数名称 | 类型 | 必选 | 示例值 | 描述 |
|---|---|---|---|---|
| file | file | 是 | - | 要转录的音频文件对象(非文件名),支持格式包括 flac、mp3、mp4、mpeg、mpga、m4a、ogg、wav、webm |
| model | string | 是 | sense-voice-small-10s-ax650 | SenseVoice 模型支持中英日粤韩等多语言自动识别 |
| language | string | 否 | - | 模型内部自动识别语言 |
| response_format | string | 否 | json | 返回格式,目前仅支持 json,默认值为 json |
Transcription(text=' Thank you. Thank you everybody. All right everybody go ahead and have a seat. How\'s everybody doing today? .....',
logprobs=None, task='transcribe', language='en', duration=334.234, segments=12, sample_rate=16000, channels=1, bit_depth=16)