Run AI on European infrastructure.

Open models through one European API. Dedicated GPUs when you need more control.

Teams building on Lyceum

References published as customers approve them.

Two product paths: inference and compute

01 / Inference

Run open models in production.

One OpenAI-compatible API. Per token, no base fee.

02 / Compute

Train and scale on European GPUs.

Virtual machines, training and clusters. Per second, no subscription.

Explore compute
03 / Our models

One API.
The right model.

All models and prices
  • Language Z.ai

    GLM-5.3

    1M context, tool use

    $1.75 in · $4.50 out per 1M tokens

  • Language Moonshot

    Kimi K3

    Long context, agentic tool use

    $3.00 in · $15.00 out per 1M tokens

  • Language Z.ai

    GLM-5.3 Flash

    Fast, low-cost, 1M context

    $0.20 in · $0.50 out per 1M tokens

  • Code DeepSeek

    DeepSeek V4 Pro

    Reasoning and coding

    $1.75 in · $3.50 out per 1M tokens

  • Code Moonshot

    Kimi K2.7 Code

    Coding-focused agentic model

    $1.25 in · $4.50 out per 1M tokens

  • Multimodal Moonshot

    Kimi K2.6

    Native multimodal, text and images

    $1.00 in · $4.00 out per 1M tokens

  • Language Qwen

    Qwen3.8 2.4T A95B

    Flagship MoE, function calling

    $2.50 in · $6.00 out per 1M tokens

  • Language Qwen

    Qwen3.8 Flash Next

    Preview of the Qwen 4 architecture

    $0.20 in · $0.50 out per 1M tokens

  • Language MiniMax

    MiniMax M3

    1M context for large documents

    $0.40 in · $2.00 out per 1M tokens

  • Language Z.ai

    GLM-5.2

    Bilingual reasoning, tool use

    $1.50 in · $4.50 out per 1M tokens

  • Language OpenAI

    gpt-oss 120B

    Open-weight 120B

    $0.15 in · $0.60 out per 1M tokens

  • Language Meta

    Llama 3.3 70B

    Instruction-following 70B

    $0.13 in · $0.40 out per 1M tokens

  • Embedding Qwen

    Qwen3 Embedding 8B

    Multilingual retrieval

    $0.01 per 1M tokens

  • Language Qwen

    Qwen3.8 27B

    Dense 27B, fine-tuning ready

    $0.40 in · $2.40 out per 1M tokens

  • Language Qwen

    Qwen3.5 9B

    Compact, fast, low cost

    $0.15 in · $0.20 out per 1M tokens

  • Language Qwen

    Qwen3 235B A22B

    Large MoE, instruction following

    $0.20 in · $0.60 out per 1M tokens

  • Language Qwen

    Qwen3 30B A3B

    Efficient MoE

    $0.10 in · $0.30 out per 1M tokens

  • Language Qwen

    Qwen3 Next 80B Thinking

    Reasoning variant

    $0.15 in · $1.20 out per 1M tokens

  • Multimodal Google

    Gemma 3 27B

    Instruction-tuned 27B

    $0.10 in · $0.30 out per 1M tokens

  • Language NousResearch

    Hermes 4 405B

    Strong reasoning, long context

    $1.00 in · $3.00 out per 1M tokens

  • Language NVIDIA

    Nemotron Ultra 253B

    Ultra-large, Llama 3.1 based

    $0.60 in · $1.80 out per 1M tokens

  • Language NVIDIA

    Nemotron 3 Nano 30B

    Compact MoE, efficient

    $0.06 in · $0.24 out per 1M tokens

  • Multimodal NVIDIA

    Nemotron 3 Nano Omni

    Omni-modal reasoning for agents

    $0.06 in · $0.24 out per 1M tokens

  • Language NVIDIA

    Cosmos 3 Super Reasoner

    Multi-step reasoning

    $0.10 in · $0.30 out per 1M tokens

  • Multimodal OpenBMB

    MiniCPM-V 4.5

    Efficient vision-language

    $0.66 in · $1.11 out per 1M tokens

04 / European sovereignty

Your AI stays close.

Only European data centres

Spain, Paris and the Nordics.

Zero data retention

Prompts and outputs are never kept or trained on.

One endpoint change

OpenAI-compatible API and plain Docker containers.

Hyper-available support

The people who run the platform, on the line.

Where we run and where we are
Lyceum offices in Berlin and Zurich, data centre regions in Paris, Spain and the Nordics PARIS SPAIN NORDICS BERLIN ZURICH
Data centre region Office
05 / GPU availability

Choose the compute.

Full pricing
Reserved and dedicated capacity

From one server for one month to a dedicated cluster. Around 200 GPUs is the largest single-customer deployment running today. A 1,000-GPU deployment is in build.

Request commitment pricing

Tell us the shape of the workload. An engineer replies with reserved and dedicated options, usually within one working day.