Flash News

id Software Co-founder: LLMs Can Serve as 'Near Lossless' Compression Devices for Massive Internet Archives

id Software co-founder John Carmack pointed out that while the industry generally discourages large language models from directly copying training data, using LLMs to compress entire internet archive-level corpora with near losslessness is an attractive technical direction. He compared the current popular 'perfect compression' competitions (like the Hutter Prize, which only targets 1GB of text) with PB-level data compression, suggesting that under different conditions of scale and precision requirements, AI-driven compression methods will present completely different trade-offs.

Carmack's idea essentially views large language models as 'lossy + reconstructible' compression engines: during the training phase, the model transforms the original text into internal parameters and contextual representations, which can theoretically be used to reconstruct a version close to the original content, but without the need to strictly restore every byte. If this model is validated on tens of PB-level internet archives, it will represent a deep coupling between data compression and AI memory.

Source: Public Information

ABAB AI Insight

Carmack的观点揭示了一条被广泛忽略的“AI压缩经济学”逻辑:在极端压缩与精确可逆之间,未来很可能会出现“高保真但可压缩”的中间层。对于互联网档案、新闻历史甚至代码仓库而言,存储成本与访问效率之间的平衡,未必需要严格“bit级重建”,而只需在可接受误差下重建内容语义与可读版本。

LLM在此场景下,可被看作“语义压缩核”——它将原始文本降维为可生成的模型参数与注意力结构,在解压端通过“重新生成”恢复内容,而非“硬性解压”。这与传统压缩算法(如gzip、LZMA)从“移除冗余字节”出发完全不同,它是从“语义重构”出发,更适合处理大规模、非结构化文本与文档集合。

在历史与知识结构维度,这一思路甚至暗示了一种“数字图书馆形态演进”:未来的档案存储系统可能不再依赖“完整文件存档”,而是依赖“模型+元数据+关键校验块”结构,用压缩空间换取检索与重建效率。Carmack的提示,不仅是对压缩竞赛的重新思考,更是对整个人类数字文明存储形态的底层追问。

AI

Source

·ABAB News
·
3 min read
·121d ago
分享: