Distributional matching for vector quantization

Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study

Fang, Xianghong, Guo, Litao, Chen, Hengchao, Zhang, Yuxuan, XiaofanXia, Song, Dingjie, Liu, Yexin, Wang, Hao, Yang, Harry, Yuan, Yuan, Sun, Qiang

Paper / source record · 2025-06-18

Paper Code

Research question and approach

The success of autoregressive models largely depends on the effectiveness of vector quantization, a technique that discretizes continuous features by mapping them to the nearest code vectors within a learnable codebook. Two critical issues in existing vector quantization methods are training instability and codebook collapse. Training instability arises from the gradient discrepancy introduced by the straight-through estimator, especially in the presence of significant quantization errors, while codebook collapse occurs when only a small subset of code vectors are utilized during training.

When this work is relevant

Use this 2025 record for its theoretical and empirical treatment of distributional matching in vector quantization. The related 2026 record has a separate identifier and overlapping material.

Author-written abstract

The success of autoregressive models largely depends on the effectiveness of vector quantization, a technique that discretizes continuous features by mapping them to the nearest code vectors within a learnable codebook. Two critical issues in existing vector quantization methods are training instability and codebook collapse. Training instability arises from the gradient discrepancy introduced by the straight-through estimator, especially in the presence of significant quantization errors, while codebook collapse occurs when only a small subset of code vectors are utilized during training. A closer examination of these issues reveals that they are primarily driven by a mismatch between the distributions of the features and code vectors, leading to unrepresentative code vectors and significant data information loss during compression. To address this, we employ the Wasserstein distance to align these two distributions, achieving near 100% codebook utilization and significantly reducing the quantization error. Both empirical and theoretical analyses validate the effectiveness of the proposed approach.

Abstract source: https://arxiv.org/abs/2506.15078. Checked 2026-09-14. Bibliographic metadata uses the linked paper record or author-maintained catalog. Results, limitations and experimental settings remain defined by the original source.

Citation

BibTeX · CITATION.cff

BibTeX is preserved from the author-maintained citation repository. The CFF uses the source metadata shown on this page. Version titles or author lists can differ; choose the version you used.

@article{fang2025enhancing,
  title={Enhancing vector quantization with distributional matching: A theoretical and empirical study},
  author={Fang, Xianghong and Guo, Litao and Chen, Hengchao and Zhang, Yuxuan and Song, Dingjie and Liu, Yexin and Wang, Hao and Yang, Harry and Yuan, Yuan and Sun, Qiang and others},
  journal={arXiv preprint arXiv:2506.15078},
  year={2025}
}