The word "local" hides the load path.
DwarfStar 4, or ds4, is a C inference engine for a small set of DeepSeek, GLM, and Qwen model families. It supports Metal, CUDA, and ROCm, but the same model layout does not run on every backend. The project chooses specific GGUF layouts instead of trying to load every GGUF file.[1]
That is a useful constraint. A local-model picker should not stop at the model name. It should name the exact quantization, backend, resident memory, disk-backed data, context setting, and unsupported paths.
"Local" names the address. It does not tell you whether the model fits, streams, stalls, or survives a restart.
Memory and storage form one route.
The ds4 hardware guide starts Apple Silicon at 64 GB for a Qwen 3.8 Q2 path with 41.73 GiB of resident weights and a 95.37 GiB n-gram table on disk. It names 96 to 128 GB as the practical DeepSeek V4 Flash class. At 128 GB, the guide also lists a resident GLM 5.3 Flash Q2 path and a streamed DeepSeek V4.1 Q2 path.[2]
Those are project recommendations, not our measurements. The useful split is still clear. Resident weights consume memory. Routed experts or auxiliary tables may stream from storage. Context state adds another load. A machine can start a model and still miss the latency target for daily work.
Context changes the speed sheet.
The project benchmark page reports separate prefill and generation rates. Its M5 Max 128 GB DeepSeek V4 Flash Q2 rows list 39.4 generated tokens per second at 2,048 tokens of context and 27.6 at 65,536. The DGX Spark rows list faster prefill but slower generation for the same published model class. These are upstream measurements on named machines and settings.[3]
Do not carry one tokens-per-second number into procurement. Save machine, backend, quantization, context, prompt shape, prefill rate, generation rate, and power mode. Measure your own workload after the model starts.
Durable cache is a behavior to test.
Ds4 stores prompt-prefix state on disk and keys it by the prefix hash. Its architecture notes say tool-call mappings also persist in those cache files. That can avoid repeated prefill work after a restart, but it creates a storage contract. You need to test cache reuse, invalidation, disk growth, and deletion with the model and client you plan to use.[4]
The repository is a real artifact. We cloned revision 0aaea5a238fb41a35106a551e73c8409dfb751ac, built its CPU targets, and ran ./ds4 --help. The build produced the CLI, server, benchmark, evaluation, and agent executables. We did not download weights or run inference, so this smoke check proves source retrieval and CPU compilation only.[1]
Buy the route, not the parameter count.
- Pick one supported model layout and one backend.
- Separate resident bytes from streamed files and cache growth.
- Choose the context length before reading a speed result.
- Run one cold prompt, one warm prefix, and one restart.
- Keep output quality and task completion beside latency.
Large local models can trade recurring provider calls for hardware, storage, setup, and test work. That can be a good trade. The receipt needs more than a successful startup.