Modelplane Modelplane docs

All recipes

Every validated recipe in one table: model, size, architecture, precision, and the verified hardware. Select a row for the full recipe.

ModelSizeArchPrecisionVerified onNotes
Qwen3-8B qwen8BDenseBF16EKS L4An 8.2B dense chat model on a single NVIDIA L4.
Qwen2.5-7B qwen7BDenseAWQ INT4Vultr A16A 7B dense chat model (AWQ INT4) on a single NVIDIA A16 on Vultr.
Qwen3-Coder-480B qwen480B A35BMoEBF16 / FP8EKS H200A 480B code MoE, multi-node BF16 over EFA or single-node FP8 on SGLang.
Qwen2.5-72B qwen72BDenseAWQ INT4AKSNebius A100H100A 72B dense chat model (AWQ INT4) on a single 80 GB GPU, on AKS and Nebius.
Kimi-K2 moonshotai1T A32BMoEINT4EKS H200A 1T MoE served prefill/decode disaggregated across two H200 nodes.
Llama-3.1-8B meta-llama8BDenseBF16EKSGKE L4An 8B dense chat model on a single NVIDIA L4.
GLM-4.5-Air zai-org106B A12BMoEGGUF IQ4_XSGKE A100A 106B MoE served from a GGUF checkpoint via llama.cpp on a single A100.
Nemotron-3.5-Lightning nvidia30B A3BMoENVFP4Nebius H100An open 30B MoE with 3B active parameters served NVFP4 on a single H100 on Nebius.
Laguna-S-2.1 poolside118B A8BMoEFP8Nebius H100A 118B code MoE served FP8 on a single 8x H100 node on Nebius.