🔥 Trending on HN

DwarfStar 4 brings large local AI to high-memory computers

2 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
inference engine(inference engine)

Software that runs a trained AI model and produces its output.

quantization(quantization)

Storing model numbers with fewer bits to reduce memory use.

KV cache(K-V cache)

Saved attention state that lets the system reuse earlier work.

What happened

DwarfStar 4, or ds4, is a new way to run large AI models locally. Salvatore Sanfilippo, also known as antirez, is known for creating Redis. The project is a C-based inference engine, not a new AI model. It combines a command-line interface, local APIs, and a coding agent in one stack. Its official site and public repository list support for DeepSeek V4 and V4.1, GLM 5.x, and Qwen3.8 Flash Next.

The Hacker News submission received about 195 points and 56 comments. That number measures community attention. It does not prove that ds4 is fast, accurate, or better than other tools. The Hacker News post is separate from the project's own technical evidence.

Why local execution is difficult

Large models need substantial memory. The project describes ds4 as an answer to running models that usually use remote servers. Its target is high-memory personal hardware. That includes Apple Silicon Macs, NVIDIA systems using CUDA, and AMD systems using ROCm.

The tradeoff is narrow support. ds4 is not a general GGUF runner. It accepts model files and layouts tested by the project. This lets the developers tune the engine for a small set of models. It also means that a model outside that set may not work.

The engineering behind it

One major technique is asymmetric two-bit quantization. Quantization stores model numbers with fewer bits. ds4 compresses routed expert sections heavily. It keeps shared sections and important decision-making values at higher precision. The project says this approach can fit frontier-class open-weight models into roughly 96GB to 128GB of unified memory.

The other key idea is the KV cache. ds4 treats this attention state as durable data. It saves cache entries to an SSD and keys them by the rendered prompt prefix. A later session can reuse matching work instead of rebuilding the same prefix. The project also describes cache state surviving process and machine restarts. This design targets long conversations and coding sessions.

What the project confirms

The hardware guide lists a Qwen3.8 Q2 path for 64GB Macs. It calls 96GB to 128GB a practical range for DeepSeek V4 Flash. Some DeepSeek V4.1 configurations belong to the 512GB class.

The project's benchmark table reports 34.4 generated tokens per second on an M5 Max with 128GB. The test used DeepSeek V4 Flash Q2 and a 32,768-token context. At 65,536 tokens, the reported generation rate was 27.6 tokens per second. The benchmark used a fixed public-domain text and a short generation probe. These are project measurements, not independent leaderboard results.

What remains uncertain

Independent tests still need to measure answer quality after quantization. Fair speed comparisons also need matching models, hardware, context sizes, and settings. The repository says ds4 is changing quickly and remains beta quality. It warns that instability and regressions are possible. Model support can also change as better open-weight models appear.

What to watch next

The important follow-up is outside replication. Can other users reproduce the speed? Does quality stay useful on each supported model? Do long contexts and tool calls remain stable? ds4 presents a focused engineering path for local AI. Its real value will depend on repeatable tests, not on Hacker News popularity.

💬 ds4 is a focused local inference engine, but its comparisons are still unsettled

HN commenters see ds4 as a deliberate attempt to optimize a small set of models for powerful consumer hardware. They also point to unresolved questions about quantization, real-world speed, and novelty versus existing runners.

  • ds4 is described as a native inference engine focused on selected models such as the DeepSeek V4 family, Qwen 3.8 Flash Next, and GLM variants. Supporters say the narrow scope reduces configuration mistakes; skeptics say it resembles llama.cpp and ask what is genuinely new.
  • Quantization results are split in user self-reports. One user self-reports poor quality from a quantized dsv4 checkpoint, while another self-reports that ds4 quants beat Unsloth quants. The checkpoint, model version, and quantization level were not aligned, so no general winner can be inferred.
  • Performance claims are also user self-reports: one user reports about 22 tokens/s on an Intel Ultra 7 255H iGPU after a small patch, and another reports using a 1M context on a 128GB M5 Max with fused TQ and Qwen 3.8 Flash Next. A separate user self-reports already having 1M context on the same machine and model, so the conditions and source of the improvement are unclear.
  • A commenter describes ds4-agent as append-only: it does not rewrite message history, allowing the reusable prefix of the KV cache to remain useful. This is presented as a benefit for agent workflows, not as an independently verified benchmark.
  • Commenters want broader Intel and AMD support. Community projects mentioned include Xenolith for Intel Xe-LP systems without XMX and smaller quantized Gemma-4 models, plus ds4go, which exposes ds4 through Go bindings and tools.
  • The C-versus-Rust debate has no settled answer. Arguments for C include the author's familiarity, the abundance of C systems code in LLM training data, and a belief that LLMs write C more effectively; the Rust counterargument emphasizes low-level access, SIMD, and compile-time safety.
  • Tool-calling examples, comparable benchmark baselines, and hardware requirements remain unclear. The 64GB-versus-96GB Mac minimum is not a user performance measurement; commenters identify it as a discrepancy between the article and the GitHub documentation.

initial digest at 56 comments (revision 1). We fetched 56 comments and sampled 56 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

A tool for running big AI on your own computer

📰 Full story: DwarfStar 4 brings large local AI to high-memory computers

DwarfStar 4 helps large AI models run on local, powerful computers.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
ds4(D-S-four)

Software for running selected large AI models locally.

quantization(quantization)

A way to store AI numbers in less space.

KV cache(K-V cache)

Saved computer work that can be reused later.

💡 The gist

  • ds4 runs some large AI models on your own computer.
  • It shrinks model data and saves used work on disk.
  • Hacker News attention is not proof of better performance.

Salvatore Sanfilippo is a programmer known for creating Redis. Redis is a computer tool for storing information quickly. His project is called DwarfStar 4, or ds4.

The official site says ds4 runs selected DeepSeek, GLM, and Qwen models. It is not the model itself. It is software that helps a computer run the model.

Why does this matter? Large AI models need lots of memory. Many models usually run on remote servers. ds4 tries to run some of them on powerful computers people can own.

The project makes a careful trade. It supports only chosen models and file layouts. That smaller target lets its developers tune the software for those models. It also means another model may not work.

One trick is called quantization. It stores some model numbers with fewer bits. This reduces the space the model needs. ds4 compresses some sections more than others. It keeps important sections more precise.

Another trick is the KV cache. It saves earlier computer work on an SSD. The computer can reuse that work during the same long conversation. This can reduce repeated calculations.

The hardware guide lists a 64GB Mac path for Qwen3.8 Q2. It lists 96GB to 128GB as a practical range for other models. Some DeepSeek V4.1 versions need a 512GB-class machine.

A project test reached 34.4 generated tokens per second. It used an M5 Max with 128GB. The test came from the project, not an outside lab. It does not prove every computer will be fast.

The Hacker News post had about 195 points and 56 comments. Those numbers show attention. They do not prove that ds4 gives the best answers. More outside testing must check speed, quality, and stability. The benchmark table shows the project's conditions.

💬 ds4 narrows its target to make local AI easier to optimize

The comments contain both enthusiasm for ds4's focused design and skepticism about how much it adds beyond llama.cpp. Most speed and quality numbers are user self-reports.

  • ds4 does not try to run every model. It focuses on selected large models and powerful personal computers. Supporters say this makes tuning easier; critics say it may be too similar to llama.cpp.
  • Users disagree about compressed model files. One user self-reports poor dsv4 quantization quality, while another self-reports that ds4 quants performed better than Unsloth's. The test details are incomplete.
  • User self-reports include about 22 tokens/s on an Intel Ultra 7 255H iGPU and a 1M context on a 128GB Mac. Another user self-reports already having that 1M context on the same setup, so these are not apples-to-apples comparisons.
  • The agent tool keeps old chat history instead of rewriting it, which can let it reuse saved KV-cache calculations. Supporters say this helps long-running agent sessions.
  • People are asking for better Intel and AMD support, and community projects such as Xenolith and ds4go extend the ecosystem.
  • The C-versus-Rust choice is still debated. Tool calling, fair benchmarks, and required memory also need clearer evidence; comments note a 64GB-versus-96GB mismatch between the article and GitHub.

initial digest at 56 comments (revision 1). We fetched 56 comments and sampled 56 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

Can big AI run on your computer?

📰 Full story: DwarfStar 4 brings large local AI to high-memory computers

DwarfStar 4 helps some big AI run nearby.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
DwarfStar 4(DwarfStar Four)

A tool for running some big AI on your computer.

Hacker News(Hacker News)

A website where people notice and discuss technology.

Salvatore Sanfilippo made Redis, a computer tool.

He also made DwarfStar 4.

DwarfStar 4 helps run big AI on your computer.

Big AI often runs on faraway computers.

Some computers do not have enough room for it.

ds4 makes some model numbers smaller.

That is like packing a big bag more tightly.

It also saves some used computer work on a fast disk.

Then the computer can reuse that work later.

ds4 supports only chosen AI models.

The official project page explains its choices.

On Hacker News, it had about 195 points and 56 comments.

Those numbers show attention, not proof that it works best.

People still need to check its speed and answers.

💬 ds4 is a small tool for running a few big AI models

Some people like its focused plan. Other people wonder whether it is truly different from older tools. The evidence so far is mostly personal reports.

  • ds4 tries to run a few chosen AI models on strong home computers. It does not try to support everything.
  • People report different results after shrinking the models. One user reports worse quality, while another reports better results than Unsloth.
  • Users report about 22 tokens each second on one Intel machine and a very long, 1-million-token memory on a 128GB Mac. Another user says that long memory already worked, so the results are not a fair test yet.
  • ds4 can remember old chat calculations and reuse them. People still want Intel and AMD support, clearer speed tests, and a clearer answer about which computers and tools work best.

initial digest at 56 comments (revision 1). We fetched 56 comments and sampled 56 across the thread. These are HN users’ reports, not independently verified facts.

Sources