Skip to main content
Serverless inference offers pay-per-request pricing with no idle compute cost, the right choice for low-volume or spiky traffic where reserving GPU replicas would be wasteful. The endpoints are OpenAI-compatible. Point any OpenAI client at the base URL below, swap in a Lyceum API key and a model name, and everything else, streaming, tool calling, usage accounting, works the same.

Try it

Enter a key and pick a model to send a real request from this page. It runs the same auth, credit and model checks as any client, so a green result means your key is valid, your org has credit, and the model ID resolves.
This sends one real chat completion and bills it to your org, the same as any other request. It is a few tokens.

Endpoint and authentication

The base URL for serverless inference is:
Authenticate with your Lyceum API key (lk_...) as a Bearer token. Most SDKs append the path for you, so you pass only the base URL above.
There is no legacy text /completions endpoint. All text generation goes through /chat/completions using the messages format.

Quickstart

Model IDs

GET /models is the authoritative list of IDs the serverless API accepts, and it reflects what is served right now. It also contains the -instant variants (same model, reasoning switched off, see Reasoning models) and the lyceum/simple, lyceum/complex and lyceum/reasoning routing aliases. Prices and context windows per model are on the models page. A model ID that is listed but has no backend available at the moment answers 404 model not found, the same as an unknown ID. Poll GET /models or retry with backoff before treating it as permanent.

Streaming

Set stream: true to receive tokens as they are generated. This is the recommended mode for anything user-facing or long-form, it lowers time-to-first-token and keeps the connection active for the full generation.

Timeouts and long-running requests

The platform allows a single request up to 300 seconds of generation time. Large context windows or long outputs on a non-streaming call can genuinely take minutes, so the most common source of “cut off” errors is a client-side timeout that fires before the response finishes, not the server.
For long outputs, stream the response. Streaming keeps data flowing over the connection, which avoids idle/read timeouts and gives you tokens as they arrive instead of waiting for the whole completion.
If you must run non-streaming for long generations, raise your client’s read timeout:
Some frameworks (for example litellm) apply a separate stream inactivity timeout that aborts a stream if no new token arrives within a fixed window. Under heavy load, time-to-first-token can spike, so if you stream through such a framework, raise its inactivity timeout as well as the overall request timeout.

Request limits

Function calling and tools

All chat models support function / tool calling using the standard OpenAI tools parameter: automatic tool calls, parallel calls and tool-result round-trips work on every model. tool_choice is forwarded to the backend, but not every model enforces it, see the warning below.
tool_choice: "required" and a named function ({"type": "function", "function": {"name": "..."}}) are currently not enforced on deepseek/deepseek-v4-pro-0813, deepseek/deepseek-v4-flash-0731 and z-ai/glm-5.3. These models can answer in plain text with no tool_calls while still reporting finish_reason: "tool_calls". In checks on 17 September 2026, tool_choice: "required" was also not enforced on deepseek/deepseek-v4.1-flash and z-ai/glm-5.2, and only some of the time on moonshotai/kimi-k2.6 and qwen/qwen3-235b-a22b-instruct-2507. moonshotai/kimi-k3, z-ai/glm-5.3-flash, minimax/minimax-m3 and qwen/qwen3.8-27b honour both forms. If your application depends on a forced tool call, check message.tool_calls and fall back when it is empty.
A small number of models prefer to answer in text rather than call a tool when tool_choice is left on "auto". On models that enforce tool_choice, set it to "required" to force a call.

Prompt caching

Prompt caching is automatic, there is nothing to enable. When consecutive requests share an identical leading prefix (for example a fixed system prompt or a large shared context), the cached portion is reused and billed at a reduced input rate. To benefit, keep the stable part of your prompt at the front and vary only the tail (the user’s latest message). Caching is best-effort, a request may or may not hit a warm cache depending on recent traffic. Cache hits are reported in the usage object:
prompt_tokens_details is absent entirely on a cache miss rather than reporting zero, so read it defensively:

Reasoning models

Several serverless models produce an explicit reasoning trace before the final answer. The trace counts toward output tokens, so set max_tokens high enough to cover both the reasoning and the visible answer. Otherwise the whole budget can be spent on reasoning and content comes back empty or truncated.
Most models return the trace in message.reasoning_content. Some backends use message.reasoning instead (observed on z-ai/glm-5.3 and, depending on the serving path, deepseek/deepseek-v4-flash-0731). Read both fields.

Controlling reasoning

The simplest switch is the top-level OpenAI parameter reasoning_effort, which Lyceum forwards to the model backend. reasoning_effort: "none" turns reasoning off on every model that supports switching. chat_template_kwargs is forwarded unchanged as well, for clients that prefer the chat-template switch; the kwarg key differs by model family (DeepSeek and Kimi: thinking, GLM and Qwen3.8: enable_thinking). Example, disabling reasoning on a model that defaults to on:
  • reasoning_effort levels other than "none" are forwarded to the model verbatim and are not reliably graded. Treat reasoning as on/off, not a depth dial.
  • Accepted values differ per model. moonshotai/kimi-k3 accepts low, medium, high, max and none. qwen/qwen3.8-27b and qwen/qwen3.8-flash-next accept none, low, medium and xhigh, and reject high, max and minimal with 400 inference request rejected.
  • Top-level enable_thinking is not honored. Use reasoning_effort or chat_template_kwargs.

Instant variants

Clients that cap max_tokens and don’t render reasoning_content (GitHub Copilot, Cursor and similar BYOK setups) can end up with an empty visible answer because the whole budget goes into reasoning. For clients that cannot pass request parameters, use the -instant model IDs: z-ai/glm-5.2-instant, qwen/qwen3.8-27b-instant and qwen/qwen3.8-flash-next-instant. Same model, same pricing, answers directly without a reasoning phase. They are listed in GET /models. z-ai/glm-5.3-instant and z-ai/glm-5.3-flash-instant are listed too, but currently write their reasoning into content, so avoid them in these clients.

Image and PDF input

Some models accept images alongside text. Send them as OpenAI-style content parts: the content of a user message becomes an array with a text part and one image_url part per file. The URL can be a public https:// URL or an inline base64 data URL.
Text-only turns can keep using a plain string for content. Images count towards the request body limit, so downscale large photos before encoding them. The dashboard playground caps the longest edge at 1568 px, which is plenty for every model listed here.

Which models accept images

Verified on 15 September 2026 against the live endpoint, with two distinct synthetic test images sent to every chat model across three sweeps. “Yes” means the model described both images correctly every time. The dashboard playground shows a Vision chip on every model in the “Yes” rows and lets you attach, drag and drop, or paste images. Click any sent message there to get the exact request.

PDF input

moonshotai/kimi-k2.7-code and moonshotai/kimi-k2.6 also read PDFs. Send the PDF exactly like an image, as an image_url part whose URL is a data:application/pdf;base64,... data URL:
  • Every other model rejects a PDF part with 400 inference request rejected, including models that accept images.
  • The OpenAI {"type": "file", "file": {...}} content part is not supported by any backend and is rejected with 400.
  • Expect PDF requests to take longer than image requests; in our checks a one-page PDF took 10 to 17 seconds on the Kimi K2 models.
The playground shows a PDF chip on these two models and accepts PDFs in the attach dialog and drop zone.

Embeddings

Text embeddings use the same base URL and key:

Error handling

All errors use the OpenAI-compatible shape with error.message, error.type and, where it applies, error.param:
Lyceum does not pass upstream error text through. Errors coming from a model backend are mapped to a small set of stable messages, so match on error.message and the HTTP status: The same shapes are used for errors that occur mid-stream. Retry on 429, 502, 503 and 504 with backoff; treat 4xx rejections as permanent for that request.

Support references

Every response carries the headers x-cloud-trace-context and traceparent. Both contain the same 32-character trace ID (the first segment of either header). Quote that trace ID together with the timestamp in UTC when you contact support about a specific request.

Serverless vs dedicated

Models and pricing

Model IDs, context windows and per-token prices for every serverless model.