Tuning Henry's Output Limits
A one-commit day adjusting token limits on the local model personas behind the churner.
A one-commit day adjusting token limits on the local model personas behind the churner.
A late night stabilizing a fleet of model profiles, capping inference hangs, and promoting a faster quantized fleet.
A single day of rapid iteration took the Gemma 4 port from a baseline to validated QKV fusion, invalidating several assumptions along the way.
Fusing QKV, inlining norm-weight application, and pushing block loads through the Q4_K path.
The standalone backend finally produced inference after fixing ABI layout, shared-memory buffers, and a simple CPU fallback.
47 commits and a progressive enablement series that ended with full local decode dispatch at 7.0 t/s and 100% correctness.
Rebrand, hybrid llama-server build, IPEX reverse-engineering notes, and a sandbox with a parallel job queue.
A 12-hour debugging session documented commit by commit until I isolated the EP bug to a fused aggregation kernel.
A day split between shipping a terminal viewer for agent sessions and chasing an expert-alignment bug in local model builds.
A long documentation-and-research push that took ArcLLM from working code to a three-phase, gated roadmap.
Turning on the Whitney venue-expert call to action, plus folding a temperature/utilization indicator into the ArcLLM stack.
One session took ArcLLM from an empty repo to streaming chat, tool calls, the responses API, persistence, and worker queues.