How BITCOS Uses AI’s Many Zeros to Go Below 1.58 Bits
ternary LLM
A language model whose weights use −1, 0, and +1.
BITCOS
A storage layout that uses the many zero values in ternary models.
bit
A small unit of computer information.
What happened
Researchers at Intel, a chip company, posted a new paper on September 14, 2026. The paper studies a smaller way to store ternary LLMs. A ternary LLM is a large language model whose weights use only −1, 0, and +1. The paper calls its method BITCOS.
The title says it breaks the 1.58-bit barrier. That phrase needs care. The 1.585 figure comes from log₂3. It assumes that three symbols appear equally often. The paper argues that real ternary models do not follow that assumption.
The background
A weight is a small number used inside an AI model. Fewer bits per weight can reduce memory use. It can also reduce the data moved during inference, when the model generates an answer.
Common ternary storage packs five ternary values into one byte. This approaches 1.585 bits per weight. Real blocks often have 128 weights. Since 128 is not divisible by five, the practical rate becomes 1.625 bits per weight.
The authors measured 29 ternary model checkpoints. Zero values made up between 29.7% and 51.5% of their weights. That imbalance creates an opportunity. A zero needs no positive or negative sign.
How BITCOS works
BITCOS stores two streams. The first is a presence bitmap. It marks whether each weight is nonzero. The second is a compacted sign stream. It stores one sign bit only for each nonzero weight.
If the zero fraction is z, the layout uses 1 + 1 − z, or 2 − z, bits per weight. At 51.5% zeros, the calculation is 2 − 0.515 = 1.485 bits. BITCOS used less storage than five-trit packing on 26 of the 29 tested models.
This does not retrain a model. It does not change the model’s values. It changes how the values are stored and unpacked. The paper describes optimized kernels for AVX-512, AVX2, and Intel Xe2 GPUs.
What the experiments show
The authors integrated BITCOS into vLLM and tested seven checkpoints. They used five platforms. These included a 64-core server CPU, a 24-core client CPU, an eight-core client CPU, an integrated GPU, and a discrete GPU.
The method helped when memory traffic limited speed. End-to-end decode improved by 1.10 to 1.18 times on the 64-core CPU. It improved by 1.02 to 1.15 times on the 24-core CPU. Gains reached 1.09 to 1.22 times on the integrated GPU. They reached 1.02 to 1.27 times on the discrete GPU.
The result was not universal. On the eight-core Lunar Lake CPU, BITCOS was slower than the two-bit baseline. The smaller data stream did not cover the extra unpacking work. This shows that smaller storage does not always mean faster inference.
What is confirmed, and what is not
The paper reports measured zero rates, storage calculations, kernel tests, and end-to-end tests. It also drew attention on Hacker News, a technology news forum. Hacker News points and comments measure community attention. They do not prove that the paper is correct.
The paper is an arXiv preprint, not proof of a finished industry standard. Independent teams still need to test other hardware and software. The paper focuses on storage and inference speed. It does not establish better answer quality. Quality still depends on the original ternary model.
What to watch next
The next questions are practical. Can other chip makers reproduce the gains? Will common runtimes support BITCOS? Does it reduce power use in real devices? The answer may decide whether this remains a clever layout or becomes part of everyday on-device AI.
Sources: the arXiv paper and the Hacker News discussion.
A New Way to Store Some AI Models in Less Space
📰 Full story: How BITCOS Uses AI’s Many Zeros to Go Below 1.58 Bits
Intel researchers found a way to use many zero values inside small AI models.
ternary LLM
A text-making AI that uses three number values.
BITCOS
A way to store ternary AI numbers in less space.
bit
A tiny unit of computer information.
💡 The gist
- Some AI numbers can fit into less space.
- Many zero values make this possible.
- The speed gain depends on the machine.
Intel, a chip company, posted the paper on arXiv. The paper studies a ternary LLM. This is a text-making AI that uses three values. Its weights are −1, 0, or +1.
Weights are tiny numbers inside an AI. The AI uses them to build answers. Smaller weights can mean smaller model files. They can also mean less data moving through memory.
Older storage methods treat the three values almost equally. That gives about 1.58 bits per weight. A bit is a tiny unit of computer information.
The researchers checked 29 ternary model checkpoints. Some had zeros in more than half their weights. The highest zero rate was 51.5%.
They created BITCOS. It uses one mark for every weight. This mark says whether the weight is zero. It uses another mark only when the weight is not zero. That second mark says positive or negative.
At 51.5% zeros, the paper gives this calculation: 2 − 0.515 = 1.485 bits. BITCOS used less space than five-trit packing on 26 of 29 models.
The researchers tested seven checkpoints. They used five hardware platforms. These included server and client CPUs. They also used an integrated GPU and a separate GPU.
The method helped some platforms. It raised final decoding speed by up to 1.18 times on one CPU. It raised speed by up to 1.27 times on one GPU. Less memory traffic helped those machines.
But one eight-core CPU became slower. It spent too much time unpacking the smaller data. So smaller storage does not always mean faster answers.
The paper also became a topic on Hacker News, a technology forum. Its reactions show attention. They do not prove the research is correct.
The paper is still an arXiv preprint. Other teams must test it. They must use different chips and software. The paper also does not show better answer quality. It mainly measures storage and speed.
Sources: the arXiv paper and Hacker News.
💬 In simple terms: how small and fast can ternary LLMs become?
The key point in the comments is not only “this may make models smaller,” but also “the speed and accuracy trade-offs are still debated.” The numbers and performance claims are commenters’ explanations, not established results here.
- Using only three weight values, −1, 0, and +1, has a theoretical baseline of about 1.58 bits for each three-way choice.
- One commenter said that if about 51% of weights are zero, smarter packing could reach 1.48 bits per weight. That is an unverified claim from the discussion.
- A machine that uses ternary weights directly can treat each weight as addition, subtraction, or no operation. Commenters think this could help CPUs, small devices, and ASICs, and that reading less from memory could help speed.
- Others argue that for post-training quantization, vector or trellis quantization and efficient GEMM kernels may work just as well. If the model is expanded back to FP16, the main benefit may be a smaller file or network transfer.
- QAT, which trains with low precision in mind, may help preserve accuracy, but training can become harder and lossless results are not guaranteed. A commenter said small-block dynamic FP4 is still not lossless on every test.
- Even four storage bits do not guarantee that four bits of useful model information survive; outliers and block scales matter. It is also unsettled whether packed weights must be expanded in memory, and whether reading less data still helps when memory bandwidth is the limit.
initial digest at 21 comments (revision 1). We fetched 21 comments and sampled 21 across the thread. These are HN users’ reports, not independently verified facts.
A Smaller Box for Some AI Numbers
📰 Full story: How BITCOS Uses AI’s Many Zeros to Go Below 1.58 Bits
Intel researchers found a way to store some AI numbers in less room.
ternary LLM
A text-making AI with three number choices.
BITCOS
A smaller way to store AI numbers.
weight
A tiny number used inside an AI.
Intel, a company that makes computer chips, studied a new idea.
A ternary LLM is a text-making AI. It uses three number choices.
The choices are minus one, zero, and plus one.
A weight is a tiny number inside the AI. The AI uses weights to make answers.
The researchers looked at 29 AI models. Some models had many zeros.
They made BITCOS. BITCOS is a smaller way to store the numbers.
It keeps a mark for each number. It adds a sign mark only when needed.
That saves room when a number is zero.
Some machines made answers faster with BITCOS. One machine made answers slower.
That machine needed extra work to read the smaller box.
The paper appeared on Hacker News, a place for computer news. Many readers noticed it there.
Many readers do not prove that a study is right.
The study is still new. Other teams must test it.
They must use other machines too.
Sources: the paper and Hacker News.
💬 For a five-year-old: a tiny AI made from three marks
People were asking whether an AI can become smaller and still stay clever. The numbers and speed ideas come from commenters’ explanations and guesses.
- The AI can give each weight one of three marks: minus, zero, or plus. Three choices take about 1.58 ordinary bits to describe.
- One person said that because about 51% of the marks are zero, they might be packed more tightly into 1.48 bits each. That has not been confirmed here.
- If a machine uses the three marks directly, each one can mean “add,” “take away,” or “do nothing.” It might help a small machine run faster, but another packing method might work just as well.
- Making the AI smaller may also make it lose useful details. Special training might help, but unusual weights can still cause trouble. The packed marks may need to be opened inside the machine, and people are still discussing whether reading less data makes it faster.
initial digest at 21 comments (revision 1). We fetched 21 comments and sampled 21 across the thread. These are HN users’ reports, not independently verified facts.
💬 HN discussion of breaking the 1.58-bit barrier for ternary LLMs
The comments discuss the log2(3) baseline for ternary weights, exploiting an excess of zeros, possible hardware and memory-bandwidth gains, and the difficulty of retaining accuracy. All numerical, performance, and losslessness claims below are commenters’ explanations, concerns, or expectations, not independently verified here.
initial digest at 21 comments (revision 1). We fetched 21 comments and sampled 21 across the thread. These are HN users’ reports, not independently verified facts.