Code for the paper "Examining Post-Training Quantization for Mixture-of-Experts: A Benchmark. Our codebase is built upon AutoGPTQ. Large Language Models (LLMs) have become foundational in the realm of ...
LLaMA-MoE is a series of open-sourced Mixture-of-Expert (MoE) models based on LLaMA and SlimPajama. We build LLaMA-MoE with the following two steps: ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results