GGUF — universal local 형식
GGUF (GPT-Generated Unified Format)는 llama.cpp 프로젝트에서 나온 single-file 형식이야. .gguf 파일 하나에 weight, tokenizer, metadata가 다 들어가. Local inference de facto 표준인 이유:
- CUDA (NVIDIA), ROCm/Vulkan (AMD), Metal (Apple Silicon), CPU AVX path 전부 지원.
- 40+ 모델 아키텍처 지원 (Llama, Qwen, Gemma, Mistral, Phi, DeepSeek 등).
- Open 모델 대부분 공개 하루 안에 community GGUF 나와.
- Ollama 내부도 모델을 GGUF blob으로 저장해.
MLX — Apple native 형식
MLX는 Apple의 머신러닝 framework. 모델은 safetensors 파일들 + config.json 디렉토리로 저장돼. Quantization은 fine group quantization을 써서 64 weight마다 scale/bias 하나를 공유하고, kernel은 Apple GPU 전용으로 짜여 있어.
- Apple Silicon 전용 (NVIDIA/AMD path 없음).
- HuggingFace의
mlx-communityorg에 미리 변환된 모델이 3,000개 넘게 올라와 있어. - Apple 하드웨어에선 decode throughput은 MLX가 앞서고, prefill latency는 GGUF가 앞서.
- Ollama v0.19+부터는 Apple Silicon에서 MLX 백엔드를 써 (preview). 그래서 사용자 입장에선 형식 고르는 일의 의미가 그만큼 작아졌어.
다른 형식들
- Safetensors — full-precision weight를 위한 HuggingFace 표준. Inference engine이 safetensors를 GGUF나 MLX로 변환해서 local에서 써.
- GGML — GGUF의 전신이야. 2026에 GGML 파일은 받지 마, 이미 deprecated됐어.
- ONNX — cross-framework runtime 형식. Classical ML에선 흔하고, LLM local-inference 세계에선 잘 안 보여.