What can an offline LLM running on a single DGX Spark actually do?
I wanted to test it with something bigger than a toy project. So I ran NVIDIA Qwen3.8-Flash-Next-NVFP4 locally through vLLM, connected it to Codex, and asked it to reimplement the core connectivity part of the NATS server in Rust.
Read post