Intel LLM-Scaler Ready With Muse Glimmer Support, Other LLMs & Features

Intel LLM-Scaler Ready With Muse Glimmer Support, Other LLMs & Features

Intel's LLM-Scaler project that was born out of their Project Battlematrix initiative aims to make it easier to run generative AI on Arc (Pro) B-Series graphics cards with the likes of vLLM, ComfyUI, SGLang, and other popular AI software in this Docker-based pre-configured AI stack. This week new LLM-Scaler releases brought same-day support for new models and other enhancements. Most notable with the Intel LLM-Scaler-vLLM beta 0.21.0-b3 release on Monday was delivering same-day support for Meta's new Muse Glimmer 30B model. Muse-Glimmer-30B with FP8 online quantization is supported by LLM-Scaler-vLLM on the likes of the Arc Pro B70. The new LLM-Scaler-vLLM beta also adds suppport for DFlash for Muse-Glimmer-30B and Qwen3.6-27B. There is also better time-to-first-token performance for Gemma-4-31B and Gemma-4-26B-A4B-it. Plus various bug fixes for this updated vLLM stack for Intel graphics. See this GitHub release for those details. Released today was LLM-Scaler-Omni beta 0.2.0-b1. This new LLM-Scaler-Omni Docker container upgrades to the ComfyUI 0.31 XPU stack, adds support for MiniMax H3 local video generation, supports Wan Animate 2 on the Arc Pro B70 and B60, and expands optimized model coverage with Wan 2.2 14B T2V Turbo, LTX-2, Z-Image / Lumina, and Krea2. This update also adds managed GGUF Q4_1 support and pinned ComfyUI-GGUF-XPU integration.

Original Source

Read the full article at Phoronix →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.