Firehose

Filtered to tagged “model integrity” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

3 AUG 2026 · Cloudflare

Serving frontier models like Kimi and GLM means fighting for GPU memory. Here's how we quantize KV caches, compress model weights, and add integrity checks to serve them faster, cheaper, and safely.