Inference
OpenAI-compatible chat completions endpoint for running inference against any model in the catalog. Use this endpoint to send prompts and receive generated responses.
Endpoints
POST
Create chat completion
Sends a chat prompt to a model and returns a completion. Supports both streaming and non-streaming response modes.
Was this page helpful?