MLOps Pipelines
Guide to authoring fine-tuning and inference pipelines in KubeMetal.
Supported Workflows
- One-Click Model Downloads: HuggingFace and Ollama weight ingestion
- LoRA / QLoRA Fine-Tuning: Supervised fine-tuning on local custom datasets
- 4-bit / 8-bit Quantization: 70% memory reduction via Apple MLX quantization
- High-Speed Local Serving: OpenAI-compatible
/v1/chat/completionsendpoints