subcluster
Vision
5 articles
Articles
June 24, 2026
Qwen2.5-VL-72B: specs, benchmarks, and how to run it on Lyceum
Qwen2.5-VL-72B matches proprietary models like GPT-4o in visual reasoning and structured data extraction. Learn how to deploy this 72-billion parameter multimodal model on European infrastructure using Lyceum's OpenAI-compatible API.
June 22, 2026
Nemotron-3-Nano-Omni: specs, benchmarks, and how to run it on Lyceum
Nemotron-3-Nano-Omni replaces fragmented vision-language-audio stacks with a single perception-to-action loop. It activates 3B parameters per token while delivering state-of-the-art multimodal reasoning.
June 21, 2026
MiniCPM-V 4.5: specs, benchmarks, and how to run it on Lyceum
MiniCPM-V 4.5 scores 77.0 on OpenCompass in an efficient 8B package. With its novel 3D-Resampler, it compresses video tokens by 96x, making long-video understanding highly cost-effective.
June 2, 2026
Run Vision Language Models on GPU Cloud: VRAM & Setup Guide
Vision language models consume massive VRAM for image tokens. Learn the exact hardware requirements and deployment strategies for production VLMs.
May 31, 2026
Multimodal AI Inference on European GPUs: Compliance and Cost Optimization
Running multimodal AI inference at scale exposes the structural flaws of hyperscaler pricing and compliance models. Engineering teams require infrastructure that provides high throughput for complex data types while maintaining strict data residency.