English
English
简体中文
日本語

Chat Completions API

The chat completions endpoint accepts a list of messages and uses a language model running on the device to generate a response. The endpoint path is /v1/chat/completions.

Before making a request, install llm-openai-api, the LLM service, and a matching model package on the device. Replace the device IP and model ID in the examples with your actual configuration. The examples below use an AX650 model for AI Pyramid; for other devices, select a platform-compatible model from the model library.

Curl Call

curl http://192.168.20.186:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-xxxxxxxx" \
  -d '{
    "model": "qwen2.5-1.5B-Int4-ax650",
    "messages": [
      {"role": "developer", "content": "You are a helpful home assistant."},
      {"role": "user", "content": "Write a one-sentence bedtime story about a unicorn."}
    ]
  }'

Python Call

from openai import OpenAI
client = OpenAI(
    api_key="sk-",
    base_url="http://192.168.20.186:8000/v1"
)

completion = client.chat.completions.create(
  model="qwen2.5-0.5B-p256-ax630c",
  messages=[
    {"role": "developer", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Hello!"}
  ]
)

print(completion.choices[0].message)

Request Parameters

Parameter Name Type Required Example Value Description
messages array Yes [{"role": "user", "content": "Hello"}] Conversation history composed of multiple messages. Text, image, audio, and other modalities are supported (depending on the model).
model string Yes qwen2.5-0.5B-p256-ax630c The model ID used to generate the reply. Multiple models are supported. Refer to the Model Introduction for selection.
audio - No - Audio output is not currently supported.
function_call - No - Function calling is not currently supported.
max_tokens integer No 1024 The maximum number of tokens the model is allowed to generate. Output will be truncated if exceeded.
response_format object No "json_object" Specifies the output format of the model. Currently, only "json_object" is supported.

Response Example

ChatCompletionMessage(content='Hello! How can I assist you today?', refusal=None, role='assistant', annotations=None, audio=None, function_call=None, tool_calls=None)
Page Tools
PDF
On This Page