cluster

Model Library

Reference coverage of every model in the Inference Studio catalogue: capabilities, context windows, hosting region and per-token prices. Serves buyers evaluating a specific open-weight model.

38 articles

Articles

August 25, 2026

Running GLM 5.1, 5.2 and 5.2 Instant in Europe: Self-Hosting and Serverless Options

Z.ai's GLM-5 series introduces 1M-token contexts and powerful agentic capabilities via a 744B MoE architecture. For European teams, running these models locally requires massive GPU clusters, making a managed serverless endpoint a highly practical alternative.

August 20, 2026

DeepSeek-V4-Flash: specs, benchmarks, and how to run it

DeepSeek-V4-Flash is a 284-billion parameter MoE model offering agentic reasoning across a 1-million token context window. Lyceum serves it via an OpenAI-compatible API from eu-north1 in the European Union, optimized for enterprise inference at $0.15 per million input tokens.

August 14, 2026

DeepSeek V4 Pro API: EU Hosting, Pricing and Context Limits

DeepSeek V4 Pro API runs in European data centres with 1M token context, $1.75/$3.50 pricing per 1M tokens, zero data retention, and full OpenAI SDK compatibility.

August 12, 2026

Where to Run Kimi Models in Europe: K2.6, K2.7 Code and K3

Moonshot AI's Kimi models deliver frontier capabilities for agentic coding. K2.6 and K2.7 Code offer 1T-parameter scale with 256K context, while K3 pushes to 2.8T parameters and a 1M-token window. European teams can run them via EU-hosted APIs to maintain data residency.

July 31, 2026

DeepSeek V4 Flash: 1M-Token Context for AI Products

DeepSeek V4 Flash introduces a 284B parameter MoE architecture with 13B active parameters, delivering low time-to-first-token latency and a 1,048,576-token context window. For AI-native products, this means high-throughput agent loops and long-context retrieval hosted natively in Europe

July 31, 2026

Kimi K3 API: Where to Run It, and What a 1M-Token Context Costs

Kimi K3 offers a 1M-token context window at $3.00 input and $15.00 output per million tokens. Deploying it on Lyceum in Europe provides GDPR compliance, zero data retention, and prompt caching at $0.75 per million tokens.

June 27, 2026

GLM-5.2: specs, benchmarks, and how to run it on Lyceum

GLM-5.2 delivers a solid 1M-token context and frontier-level coding performance at a fraction of the cost. Deploy it on European infrastructure via our Serverless Inference API.

June 27, 2026

Qwen3-Embedding-8B: specs, benchmarks, and how to run it on Lyceum

Qwen3-Embedding-8B delivers state-of-the-art retrieval performance across 100+ languages. Built on the Qwen3 foundation, it supports customizable output dimensions and instruction-aware queries for complex RAG pipelines.

June 27, 2026

Wan Image: specs, benchmarks, and how to run it on Lyceum

Wan Image delivers photorealistic generation with advanced prompt adherence. Here is how to deploy it on Lyceum Technology.

June 26, 2026

Qwen3-32B: specs, benchmarks, and how to run it on Lyceum

Qwen3-32B introduces a dual-mode architecture that smoothly switches between complex logical reasoning and efficient general-purpose chat. Now available on Lyceum's EU-hosted infrastructure, it offers a highly capable alternative to larger 70B+ models.

June 26, 2026

Qwen3.5-397B-A17B: specs, benchmarks, and how to run it on Lyceum

Qwen3.5-397B-A17B combines a massive 397-billion parameter knowledge base with an efficient 17B active-parameter routing. It delivers frontier-level coding and multimodal reasoning at a fraction of the compute cost.

June 25, 2026

Qwen3-235B-A22B: specs, benchmarks, and how to run it on Lyceum

Qwen3-235B-A22B-Instruct-2507 is Alibaba's flagship Mixture-of-Experts model, activating only 22B parameters per token for efficient performance. With a 256K context window and strong coding capabilities, it rivals top-tier proprietary models.

June 25, 2026

Qwen3-30B-A3B: specs, benchmarks, and how to run it on Lyceum

Qwen3-30B-A3B activates only 3 billion parameters per token, delivering the reasoning capabilities of a 30B model at high speeds. Learn how to deploy this cost-efficient MoE model on Lyceum's EU-sovereign infrastructure.

June 24, 2026

Nemotron-Ultra-253B: specs, benchmarks, and how to run it on Lyceum

Nemotron-Ultra-253B delivers frontier-level reasoning and coding capabilities while fitting on a single 8xH100 node. By using Neural Architecture Search (NAS) to compress the Llama 3.1 405B architecture, NVIDIA created a highly efficient model for complex math, RAG, and tool calling.

June 24, 2026

Qwen2.5-VL-72B: specs, benchmarks, and how to run it on Lyceum

Qwen2.5-VL-72B matches proprietary models like GPT-4o in visual reasoning and structured data extraction. Learn how to deploy this 72-billion parameter multimodal model on European infrastructure using Lyceum's OpenAI-compatible API.

June 23, 2026

Nemotron-3-Super-120b-a12b: specs, benchmarks, and how to run it on Lyceum

Nemotron-3-Super-120b-a12b delivers 120B-parameter reasoning with the inference cost of a 12B model. Built on a hybrid Mamba-Transformer architecture, it excels at multi-agent workflows and long-context tasks.

June 23, 2026

Nemotron-3-Ultra-550b: specs, benchmarks, and how to run it on Lyceum

Nemotron-3-Ultra-550b is a frontier-scale open model designed for complex reasoning, coding, and deep research. With native speculative decoding, it delivers high throughput for agentic tasks.

June 22, 2026

Nemotron-3-Nano-30B: specs, benchmarks, and how to run it on Lyceum

NVIDIA's Nemotron-3-Nano-30B-A3B combines a Mamba-Transformer architecture with a Mixture-of-Experts design to deliver top-tier reasoning at a fraction of the compute cost. Here is how to deploy it on Lyceum's EU-sovereign infrastructure.

June 22, 2026

Nemotron-3-Nano-Omni: specs, benchmarks, and how to run it on Lyceum

Nemotron-3-Nano-Omni replaces fragmented vision-language-audio stacks with a single perception-to-action loop. It activates 3B parameters per token while delivering state-of-the-art multimodal reasoning.

June 21, 2026

MiniCPM-V 4.5: specs, benchmarks, and how to run it on Lyceum

MiniCPM-V 4.5 scores 77.0 on OpenCompass in an efficient 8B package. With its novel 3D-Resampler, it compresses video tokens by 96x, making long-video understanding highly cost-effective.

June 21, 2026

MiniMax-M2.5: specs, benchmarks, and how to run it on Lyceum

MiniMax-M2.5 delivers frontier-level coding performance at a fraction of the cost of proprietary models. Learn how to deploy this 230B parameter MoE model on Lyceum's serverless platform.

June 20, 2026

Kimi-K2.6: specs, benchmarks, and how to run it on Lyceum

Kimi-K2.6 introduces a 300-agent swarm architecture and native multimodal capabilities for complex software engineering tasks. Deploy it instantly via Lyceum's OpenAI-compatible API.

June 20, 2026

Llama-3.3-70B: specs, benchmarks, and how to run it on Lyceum

Llama-3.3-70B-Instruct is a text-only refresh that delivers state-of-the-art performance in reasoning, math, and coding. It matches the capabilities of much larger models while maintaining the efficiency of a 70B parameter architecture.

June 19, 2026

Hermes-4-70B: specs, benchmarks, and how to run it on Lyceum

Hermes-4-70B introduces a hybrid reasoning mode and strict JSON schema adherence for complex logic tasks. Learn how to deploy it using Lyceum's OpenAI-compatible API with GDPR-compliant processing in European data centres.

June 19, 2026

Image Ultra: specs, benchmarks, and how to run it on Lyceum

Image Ultra delivers high-quality image generation in under one second. Designed for latency-sensitive applications, it offers a drop-in OpenAI-compatible API on EU-sovereign infrastructure.

June 18, 2026

gpt-oss-120b: specs, benchmarks, and how to run it on Lyceum

gpt-oss-120b brings OpenAI's reasoning capabilities to the open-source ecosystem. With 117B parameters and a sparse MoE architecture, it delivers o4-mini-level performance while fitting on a single 80GB GPU.

June 18, 2026

Hermes-4-405B: specs, benchmarks, and how to run it on Lyceum

Hermes-4-405B introduces a hybrid reasoning mode that balances fast responses with deep, think-tag deliberation. Now available on Lyceum's European infrastructure, it delivers strong math and coding performance without the censorship of proprietary models.

June 17, 2026

GLM-5.1: specs, benchmarks, and how to run it on Lyceum

GLM-5.1 is a Mixture-of-Experts model with 754B parameters and 40B active per token, built for sustained, multi-step software engineering tasks. With a leading SWE-Bench Pro score among the models on its own card, it offers an open-weight alternative to frontier proprietary models.

June 16, 2026

FLUX.1 Dev: specs, benchmarks, and how to run it on Lyceum

FLUX.1 Dev brings strong prompt adherence and photorealism to open-weights image generation. Learn how to deploy this 12B parameter rectified flow transformer on Lyceum's EU-hosted infrastructure.

June 16, 2026

FLUX.2 Klein: specs, benchmarks, and how to run it on Lyceum

FLUX.2 Klein optimizes the speed-to-quality ratio for AI image generation. With a unified architecture for text-to-image and editing, it delivers photorealistic 1024x1024 outputs in under a second.

June 15, 2026

Cosmos3-Super-Reasoner: specs, benchmarks, and how to run it on Lyceum

Cosmos3-Super-Reasoner is the 32B reasoner tower of NVIDIA's Cosmos 3 Super, built for physical AI, robotics, and complex video understanding. It takes text, images, and video to reason about real-world environments.

June 15, 2026

DeepSeek-V4-Pro: specs, benchmarks, and how to run it on Lyceum

DeepSeek-V4-Pro delivers frontier-level reasoning and a massive 1M-token context window. Learn how to deploy it through Lyceum's OpenAI-compatible API with simple per-token pricing.

June 2, 2026

Run Vision Language Models on GPU Cloud: VRAM & Setup Guide

Vision language models consume massive VRAM for image tokens. Learn the exact hardware requirements and deployment strategies for production VLMs.

May 31, 2026

Multimodal AI Inference on European GPUs: Compliance and Cost Optimization

Running multimodal AI inference at scale exposes the structural flaws of hyperscaler pricing and compliance models. Engineering teams require infrastructure that provides high throughput for complex data types while maintaining strict data residency.

May 29, 2026

Deploying Microsoft Phi-4 Inference on GPU Cloud: A Production Guide

Microsoft's Phi-4 delivers advanced reasoning at a fraction of the size of frontier models. Moving from local testing to production inference requires strict memory management and the right infrastructure stack.

May 29, 2026

Deploy Qwen 2.5 72B on GPU Cloud: VRAM Sizing and vLLM Setup

Running Qwen 2.5 72B in production requires strict memory management and the right infrastructure. Learn how to calculate VRAM requirements, configure vLLM, and deploy on EU-sovereign GPUs without hyperscaler price premiums.

May 28, 2026

Deploy Gemma 3 on European GPU Cloud: VRAM, Setup, and GDPR Compliance

Google's Gemma 3 models bring multimodal capabilities and 128K context windows to open weights AI. Running them in production requires careful VRAM planning and infrastructure that guarantees data residency.

May 27, 2026

Deploy DeepSeek R1 on European GPU Cloud: VRAM, Costs, and Compliance

Deploying DeepSeek R1 requires massive VRAM and strict data governance. Learn how to size your hardware and run production inference on EU-sovereign infrastructure without hyperscaler markups.