语音合成接口接收文本,并使用设备上运行的语音合成模型生成音频文件。本文示例使用适用于 AX630C 平台的 MeloTTS 模型包。
调用接口前,需要在设备上安装语音合成模型和示例所需的工具。模型支持情况请参考模型库。
llm-model-melotts-en-us 模型包。apt install llm-model-melotts-en-us ffmpeg 工具。apt install ffmpeg systemctl restart llm-openai-api 以下示例在 PC 端使用 OpenAI Python SDK 提交文本并接收生成的音频。调用远程设备时,请将 base_url 中的地址替换为设备的实际 IP 地址。
from pathlib import Path
from openai import OpenAI
client = OpenAI(
api_key="sk-",
base_url="http://192.168.20.186:8000/v1"
)
speech_file_path = Path(__file__).parent / "speech.mp3"
with client.audio.speech.with_streaming_response.create(
model="melotts-en-us",
voice="alloy",
input="The quick brown fox jumped over the lazy dog."
) as response:
response.stream_to_file(speech_file_path) | 参数名称 | 类型 | 必选 | 示例值 | 描述 |
|---|---|---|---|---|
| input | string | 是 | "你好,欢迎使用系统" | 要生成音频的文本内容,最大长度为 1024 个字符 |
| model | string | 是 | melotts-zh-cn | 可用的 TTS 模型,包括 melotts-ja-jp、melotts-zh-cn 和 melotts-en-us 等 |
| voice | - | 否 | - | MeloTTS 模型不支持语音风格选择 |
| response_format | string | 否 | mp3 | 音频输出格式,支持 mp3, opus, aac, flac, wav, pcm 等 |
| speed | number | 否 | 1.0 | 生成语音的速度,范围为 0.25 ~ 2.0,默认值为 1.0 |
speech_file_path 路径下。