Try it
Enter a key and pick a model to send a real request from this page. It runs the same auth, credit and model checks as any client, so a green result means your key is valid, your org has credit, and the model ID resolves.This sends one real chat completion and bills it to your org, the same as any other request. It is a few tokens.
Endpoint and authentication
The base URL for serverless inference is:lk_...) as a Bearer token. Most SDKs append the path for you, so you pass only the base URL above.
There is no legacy text
/completions endpoint. All text generation goes through /chat/completions using the messages format.Quickstart
- Python
- curl
Model IDs
GET /models is the authoritative list of IDs the serverless API accepts, and it reflects what is served right now. It also contains the -instant variants (same model, reasoning switched off, see Reasoning models) and the lyceum/simple, lyceum/complex and lyceum/reasoning routing aliases. Prices and context windows per model are on the models page.
A model ID that is listed but has no backend available at the moment answers
404 model not found, the same as an unknown ID. Poll GET /models or retry with backoff before treating it as permanent.
Streaming
Setstream: true to receive tokens as they are generated. This is the recommended mode for anything user-facing or long-form, it lowers time-to-first-token and keeps the connection active for the full generation.
Timeouts and long-running requests
The platform allows a single request up to 300 seconds of generation time. Large context windows or long outputs on a non-streaming call can genuinely take minutes, so the most common source of “cut off” errors is a client-side timeout that fires before the response finishes, not the server. If you must run non-streaming for long generations, raise your client’s read timeout:Request limits
Function calling and tools
All chat models support function / tool calling using the standard OpenAItools parameter: automatic tool calls, parallel calls and tool-result round-trips work on every model. tool_choice is forwarded to the backend, but not every model enforces it, see the warning below.
A small number of models prefer to answer in text rather than call a tool when
tool_choice is left on "auto". On models that enforce tool_choice, set it to "required" to force a call.Prompt caching
Prompt caching is automatic, there is nothing to enable. When consecutive requests share an identical leading prefix (for example a fixed system prompt or a large shared context), the cached portion is reused and billed at a reduced input rate. To benefit, keep the stable part of your prompt at the front and vary only the tail (the user’s latest message). Caching is best-effort, a request may or may not hit a warm cache depending on recent traffic. Cache hits are reported in the usage object:prompt_tokens_details is absent entirely on a cache miss rather than reporting zero, so read it defensively:Reasoning models
Several serverless models produce an explicit reasoning trace before the final answer. The trace counts toward output tokens, so setmax_tokens high enough to cover both the reasoning and the visible answer. Otherwise the whole budget can be spent on reasoning and content comes back empty or truncated.
Most models return the trace in
message.reasoning_content. Some backends use message.reasoning instead (observed on z-ai/glm-5.3 and, depending on the serving path, deepseek/deepseek-v4-flash-0731). Read both fields.Controlling reasoning
The simplest switch is the top-level OpenAI parameterreasoning_effort, which Lyceum forwards to the model backend. reasoning_effort: "none" turns reasoning off on every model that supports switching. chat_template_kwargs is forwarded unchanged as well, for clients that prefer the chat-template switch; the kwarg key differs by model family (DeepSeek and Kimi: thinking, GLM and Qwen3.8: enable_thinking).
Example, disabling reasoning on a model that defaults to on:
reasoning_effortlevels other than"none"are forwarded to the model verbatim and are not reliably graded. Treat reasoning as on/off, not a depth dial.- Accepted values differ per model.
moonshotai/kimi-k3acceptslow,medium,high,maxandnone.qwen/qwen3.8-27bandqwen/qwen3.8-flash-nextacceptnone,low,mediumandxhigh, and rejecthigh,maxandminimalwith400 inference request rejected. - Top-level
enable_thinkingis not honored. Usereasoning_effortorchat_template_kwargs.
Instant variants
Clients that capmax_tokens and don’t render reasoning_content (GitHub Copilot, Cursor and similar BYOK setups) can end up with an empty visible answer because the whole budget goes into reasoning. For clients that cannot pass request parameters, use the -instant model IDs: z-ai/glm-5.2-instant, qwen/qwen3.8-27b-instant and qwen/qwen3.8-flash-next-instant. Same model, same pricing, answers directly without a reasoning phase. They are listed in GET /models. z-ai/glm-5.3-instant and z-ai/glm-5.3-flash-instant are listed too, but currently write their reasoning into content, so avoid them in these clients.
Image and PDF input
Some models accept images alongside text. Send them as OpenAI-style content parts: thecontent of a user message becomes an array with a text part and one image_url part per file. The URL can be a public https:// URL or an inline base64 data URL.
content. Images count towards the request body limit, so downscale large photos before encoding them. The dashboard playground caps the longest edge at 1568 px, which is plenty for every model listed here.
Which models accept images
Verified on 15 September 2026 against the live endpoint, with two distinct synthetic test images sent to every chat model across three sweeps. “Yes” means the model described both images correctly every time.
The dashboard playground shows a Vision chip on every model in the “Yes” rows and lets you attach, drag and drop, or paste images. Click any sent message there to get the exact request.
PDF input
moonshotai/kimi-k2.7-code and moonshotai/kimi-k2.6 also read PDFs. Send the PDF exactly like an image, as an image_url part whose URL is a data:application/pdf;base64,... data URL:
- Every other model rejects a PDF part with
400 inference request rejected, including models that accept images. - The OpenAI
{"type": "file", "file": {...}}content part is not supported by any backend and is rejected with400. - Expect PDF requests to take longer than image requests; in our checks a one-page PDF took 10 to 17 seconds on the Kimi K2 models.
Embeddings
Text embeddings use the same base URL and key:Error handling
All errors use the OpenAI-compatible shape witherror.message, error.type and, where it applies, error.param:
error.message and the HTTP status:
The same shapes are used for errors that occur mid-stream. Retry on
429, 502, 503 and 504 with backoff; treat 4xx rejections as permanent for that request.
Support references
Every response carries the headersx-cloud-trace-context and traceparent. Both contain the same 32-character trace ID (the first segment of either header). Quote that trace ID together with the timestamp in UTC when you contact support about a specific request.
Serverless vs dedicated
Models and pricing
Model IDs, context windows and per-token prices for every serverless model.