Lyceum
Inference
  • Serverless inference
  • Dedicated endpoints
  • Model library
  • Inference pricing
  • API documentation
Compute
  • GPU virtual machines
  • Serverless training
  • Dedicated GPU clusters
  • Compute pricing
Pricing Docs
Company
  • About
  • Team
  • Careers
  • Trust centre
  • Magazine
Log in Start building Welcome back, Open dashboard
Inference
  • Serverless inference
  • Dedicated endpoints
  • Model library
  • Inference pricing
  • API documentation
Compute
  • GPU virtual machines
  • Serverless training
  • Dedicated GPU clusters
  • Compute pricing
Pricing Docs
Company
  • About
  • Team
  • Careers
  • Trust centre
  • Magazine
Start building Log in Open dashboard
  1. Home
  2. › Magazine
  3. › Inference
  4. › Inference Serving
  5. › Memory

subcluster

Memory

4 articles

Articles

June 4, 2026

Long Context Inference: GPU Requirements & VRAM Guide

Context kills VRAM. Learn the exact math behind KV cache bottlenecks and how to architect your GPU infrastructure for 128K+ token workloads.

May 31, 2026

LLM Context Length vs. GPU Memory: Calculating VRAM Requirements

Parameter count only tells half the story. Learn how to calculate the exact GPU memory required for long-context LLM inference and avoid catastrophic Out-of-Memory errors in production.

May 18, 2026

GGUF vs GPTQ vs AWQ: The Definitive LLM Quantization Framework

We break down the exact performance, memory, and throughput differences between GGUF, GPTQ, and AWQ for production inference.

February 23, 2026

KV Cache Memory Calculation for LLMs: A Technical Guide

Calculating KV cache memory is critical for preventing Out-of-Memory errors and optimizing throughput in LLM deployments. This guide breaks down the mathematical formulas and architectural variables that determine your GPU memory footprint.

Other subclusters

Endpoint Types 12 Autoscaling 2 Cold Starts 2 Throughput 7 Multi-Model 2 Routing 1
← All articles
European AI infrastructure.
Berlin and Zurich.
Live status
  • Docs
  • Models
  • Pricing
  • Trust
  • Careers
  • Contact

© 2026 Lyceum Technology Germany GmbH

  • Privacy
  • Terms
  • Imprint

Get started with GPU compute in minutes

Book a Demo

Cookies

We use cookies to measure traffic. Rejecting keeps everything working. Learn more