<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Vllm on kernelsparks</title><link>https://kernelsparks.com/tags/vllm/</link><description>Recent content in Vllm on kernelsparks</description><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 22 Sep 2026 00:00:00 +0300</lastBuildDate><atom:link href="https://kernelsparks.com/tags/vllm/index.xml" rel="self" type="application/rss+xml"/><item><title>Choosing an Agentic Model for a 2× DGX Spark Cluster (September 2026)</title><link>https://kernelsparks.com/posts/two-dgx-spark-model-choice-sept-2026/</link><pubDate>Tue, 22 Sep 2026 00:00:00 +0300</pubDate><guid>https://kernelsparks.com/posts/two-dgx-spark-model-choice-sept-2026/</guid><description>DeepSeek-V4.1-Flash at 552B, GLM-5.3-Flash at 320B, Qwen3.8-Flash-Next at 125B: which one goes on a pair of DGX Sparks behind Claude Code. Memory arithmetic decides before any benchmark does, 6B active parameters do not buy 90 tok/s, and for an agent the deciding factor is one vLLM flag, not decode speed.</description></item></channel></rss>