<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Vladislav Troinich — Lead Go Engineer &amp; Software Architect</title><link>http://troinich.pro/</link><description>Recent content on Vladislav Troinich — Lead Go Engineer &amp; Software Architect</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 25 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="http://troinich.pro/index.xml" rel="self" type="application/rss+xml"/><item><title>What can an offline LLM running on a single DGX Spark actually do?</title><link>http://troinich.pro/posts/self-hosted-llm-nats-rust/</link><pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate><guid>http://troinich.pro/posts/self-hosted-llm-nats-rust/</guid><description>&lt;h1 id="what-can-an-offline-llm-do-on-a-single-dgx-spark-rewriting-nats-core-in-rust-with-nvidia-qwen38-flash-next-nvfp4">What Can an Offline LLM Do on a Single DGX Spark? Rewriting NATS Core in Rust with NVIDIA Qwen3.8-Flash-Next-NVFP4&lt;/h1>
&lt;h2 id="intro">Intro&lt;/h2>
&lt;p>When you use frontier LLMs that are served somewhere in a data center, you can&amp;rsquo;t be sure that your work will stay private. Big companies claim that they are not using your chats and data, but can you be sure? Especially after this month&amp;rsquo;s scandal involving OpenAI and private math research.
What if you need to be certain that your work won&amp;rsquo;t leak? You don&amp;rsquo;t have many options other than running models locally. Yes, they will be slower and not as smart as frontier models (of course, I&amp;rsquo;m talking about a consumer-grade setup).
I have an ASUS Ascent GX10 (similar to an NVIDIA DGX Spark) for local LLMs. I&amp;rsquo;m keen on running local LLMs on it and interested in what I can do with them.
I was curious: what substantial piece of work could I delegate to an AI agent and get done?
I&amp;rsquo;ve used &lt;a href="https://nats.io">NATS&lt;/a> at work for a long time.
I wondered: what if I asked an agent to rewrite the NATS server (only the core connectivity part) in Rust? There are many stories of development teams rewriting their projects in Rust and having them magically run faster. I knew that NATS was a carefully engineered project, and I didn&amp;rsquo;t expect it to become faster in Rust. But I thought it would be an interesting exercise for AI, as well as easy to measure and test. Below, I&amp;rsquo;ll describe the tools I used and the results of this experiment.&lt;/p></description></item></channel></rss>