GGUF (量化版): http://t.cn/AX0GMA5v…
原始模型:
引用:We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well !
Meet MIX-STQ1_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower x.com/TencentHunyuan…