🔥 Trending on HN

How BITCOS Uses AI’s Many Zeros to Go Below 1.58 Bits

3 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
ternary LLM

A language model whose weights use −1, 0, and +1.

BITCOS

A storage layout that uses the many zero values in ternary models.

bit

A small unit of computer information.

What happened

Researchers at Intel, a chip company, posted a new paper on September 14, 2026. The paper studies a smaller way to store ternary LLMs. A ternary LLM is a large language model whose weights use only −1, 0, and +1. The paper calls its method BITCOS.

The title says it breaks the 1.58-bit barrier. That phrase needs care. The 1.585 figure comes from log₂3. It assumes that three symbols appear equally often. The paper argues that real ternary models do not follow that assumption.

The background

A weight is a small number used inside an AI model. Fewer bits per weight can reduce memory use. It can also reduce the data moved during inference, when the model generates an answer.

Common ternary storage packs five ternary values into one byte. This approaches 1.585 bits per weight. Real blocks often have 128 weights. Since 128 is not divisible by five, the practical rate becomes 1.625 bits per weight.

The authors measured 29 ternary model checkpoints. Zero values made up between 29.7% and 51.5% of their weights. That imbalance creates an opportunity. A zero needs no positive or negative sign.

How BITCOS works

BITCOS stores two streams. The first is a presence bitmap. It marks whether each weight is nonzero. The second is a compacted sign stream. It stores one sign bit only for each nonzero weight.

If the zero fraction is z, the layout uses 1 + 1 − z, or 2 − z, bits per weight. At 51.5% zeros, the calculation is 2 − 0.515 = 1.485 bits. BITCOS used less storage than five-trit packing on 26 of the 29 tested models.

This does not retrain a model. It does not change the model’s values. It changes how the values are stored and unpacked. The paper describes optimized kernels for AVX-512, AVX2, and Intel Xe2 GPUs.

What the experiments show

The authors integrated BITCOS into vLLM and tested seven checkpoints. They used five platforms. These included a 64-core server CPU, a 24-core client CPU, an eight-core client CPU, an integrated GPU, and a discrete GPU.

The method helped when memory traffic limited speed. End-to-end decode improved by 1.10 to 1.18 times on the 64-core CPU. It improved by 1.02 to 1.15 times on the 24-core CPU. Gains reached 1.09 to 1.22 times on the integrated GPU. They reached 1.02 to 1.27 times on the discrete GPU.

The result was not universal. On the eight-core Lunar Lake CPU, BITCOS was slower than the two-bit baseline. The smaller data stream did not cover the extra unpacking work. This shows that smaller storage does not always mean faster inference.

What is confirmed, and what is not

The paper reports measured zero rates, storage calculations, kernel tests, and end-to-end tests. It also drew attention on Hacker News, a technology news forum. Hacker News points and comments measure community attention. They do not prove that the paper is correct.

The paper is an arXiv preprint, not proof of a finished industry standard. Independent teams still need to test other hardware and software. The paper focuses on storage and inference speed. It does not establish better answer quality. Quality still depends on the original ternary model.

What to watch next

The next questions are practical. Can other chip makers reproduce the gains? Will common runtimes support BITCOS? Does it reduce power use in real devices? The answer may decide whether this remains a clever layout or becomes part of everyday on-device AI.

Sources: the arXiv paper and the Hacker News discussion.

💬 HN discussion of breaking the 1.58-bit barrier for ternary LLMs

The comments discuss the log2(3) baseline for ternary weights, exploiting an excess of zeros, possible hardware and memory-bandwidth gains, and the difficulty of retaining accuracy. All numerical, performance, and losslessness claims below are commenters’ explanations, concerns, or expectations, not independently verified here.

  • If each weight has three states, −1, 0, and +1, the information content of an even three-way choice is log2(3)≈1.58 bits. Some commenters found that more meaningful than saying “one trit.”
  • One commenter explained that if real model weights are zero about 51% of the time, better coding could lower the average from 1.58 to 1.48 bits per weight. This is a claim made in the discussion, not independently checked here.
  • If hardware can operate on ternary weights directly, a weight becomes addition, subtraction, or skipping the operation, which commenters expect could help CPUs, edge devices, or ASICs. Another commenter said reading less data may help when inference is memory-bandwidth-bound. These are user expectations or self-reports.
  • The counterargument is that for PTQ, vector quantization, trellis methods, or efficient GEMM kernels such as FLUTE may be comparably good. If a codebook is only used to reconstruct an FP16 model, the main savings may be storage or network transfer rather than arithmetic.
  • Accuracy is unsettled. Commenters distinguish QAT or low-bit training from PTQ; some say QAT helps preserve accuracy, while others warn about harder or less efficient training and no guarantee of lossless results. One commenter reported that even small-block dynamic FP4 is not lossless on every benchmark.
  • Another objection is that 4 storage bits do not automatically preserve four bits of useful information: block scale and range, clipping, outliers, and which weights matter all affect the result. The discussion also leaves open whether packed data must be expanded in memory—for example, five trits per byte—and whether reduced reads still help when memory bandwidth is the bottleneck.

initial digest at 21 comments (revision 1). We fetched 21 comments and sampled 21 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

A New Way to Store Some AI Models in Less Space

📰 Full story: How BITCOS Uses AI’s Many Zeros to Go Below 1.58 Bits

Intel researchers found a way to use many zero values inside small AI models.

2 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
ternary LLM

A text-making AI that uses three number values.

BITCOS

A way to store ternary AI numbers in less space.

bit

A tiny unit of computer information.

💡 The gist

  • Some AI numbers can fit into less space.
  • Many zero values make this possible.
  • The speed gain depends on the machine.

Intel, a chip company, posted the paper on arXiv. The paper studies a ternary LLM. This is a text-making AI that uses three values. Its weights are −1, 0, or +1.

Weights are tiny numbers inside an AI. The AI uses them to build answers. Smaller weights can mean smaller model files. They can also mean less data moving through memory.

Older storage methods treat the three values almost equally. That gives about 1.58 bits per weight. A bit is a tiny unit of computer information.

The researchers checked 29 ternary model checkpoints. Some had zeros in more than half their weights. The highest zero rate was 51.5%.

They created BITCOS. It uses one mark for every weight. This mark says whether the weight is zero. It uses another mark only when the weight is not zero. That second mark says positive or negative.

At 51.5% zeros, the paper gives this calculation: 2 − 0.515 = 1.485 bits. BITCOS used less space than five-trit packing on 26 of 29 models.

The researchers tested seven checkpoints. They used five hardware platforms. These included server and client CPUs. They also used an integrated GPU and a separate GPU.

The method helped some platforms. It raised final decoding speed by up to 1.18 times on one CPU. It raised speed by up to 1.27 times on one GPU. Less memory traffic helped those machines.

But one eight-core CPU became slower. It spent too much time unpacking the smaller data. So smaller storage does not always mean faster answers.

The paper also became a topic on Hacker News, a technology forum. Its reactions show attention. They do not prove the research is correct.

The paper is still an arXiv preprint. Other teams must test it. They must use different chips and software. The paper also does not show better answer quality. It mainly measures storage and speed.

Sources: the arXiv paper and Hacker News.

💬 In simple terms: how small and fast can ternary LLMs become?

The key point in the comments is not only “this may make models smaller,” but also “the speed and accuracy trade-offs are still debated.” The numbers and performance claims are commenters’ explanations, not established results here.

  • Using only three weight values, −1, 0, and +1, has a theoretical baseline of about 1.58 bits for each three-way choice.
  • One commenter said that if about 51% of weights are zero, smarter packing could reach 1.48 bits per weight. That is an unverified claim from the discussion.
  • A machine that uses ternary weights directly can treat each weight as addition, subtraction, or no operation. Commenters think this could help CPUs, small devices, and ASICs, and that reading less from memory could help speed.
  • Others argue that for post-training quantization, vector or trellis quantization and efficient GEMM kernels may work just as well. If the model is expanded back to FP16, the main benefit may be a smaller file or network transfer.
  • QAT, which trains with low precision in mind, may help preserve accuracy, but training can become harder and lossless results are not guaranteed. A commenter said small-block dynamic FP4 is still not lossless on every test.
  • Even four storage bits do not guarantee that four bits of useful model information survive; outliers and block scales matter. It is also unsettled whether packed weights must be expanded in memory, and whether reading less data still helps when memory bandwidth is the limit.

initial digest at 21 comments (revision 1). We fetched 21 comments and sampled 21 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

A Smaller Box for Some AI Numbers

📰 Full story: How BITCOS Uses AI’s Many Zeros to Go Below 1.58 Bits

Intel researchers found a way to store some AI numbers in less room.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
ternary LLM

A text-making AI with three number choices.

BITCOS

A smaller way to store AI numbers.

weight

A tiny number used inside an AI.

Intel, a company that makes computer chips, studied a new idea.

A ternary LLM is a text-making AI. It uses three number choices.

The choices are minus one, zero, and plus one.

A weight is a tiny number inside the AI. The AI uses weights to make answers.

The researchers looked at 29 AI models. Some models had many zeros.

They made BITCOS. BITCOS is a smaller way to store the numbers.

It keeps a mark for each number. It adds a sign mark only when needed.

That saves room when a number is zero.

Some machines made answers faster with BITCOS. One machine made answers slower.

That machine needed extra work to read the smaller box.

The paper appeared on Hacker News, a place for computer news. Many readers noticed it there.

Many readers do not prove that a study is right.

The study is still new. Other teams must test it.

They must use other machines too.

Sources: the paper and Hacker News.

💬 For a five-year-old: a tiny AI made from three marks

People were asking whether an AI can become smaller and still stay clever. The numbers and speed ideas come from commenters’ explanations and guesses.

  • The AI can give each weight one of three marks: minus, zero, or plus. Three choices take about 1.58 ordinary bits to describe.
  • One person said that because about 51% of the marks are zero, they might be packed more tightly into 1.48 bits each. That has not been confirmed here.
  • If a machine uses the three marks directly, each one can mean “add,” “take away,” or “do nothing.” It might help a small machine run faster, but another packing method might work just as well.
  • Making the AI smaller may also make it lose useful details. Special training might help, but unusual weights can still cause trouble. The packed marks may need to be opened inside the machine, and people are still discussing whether reading less data makes it faster.

initial digest at 21 comments (revision 1). We fetched 21 comments and sampled 21 across the thread. These are HN users’ reports, not independently verified facts.

Sources