Fine-tuning model¶
As an analytical predictor, Kavier's predictions are based on properties of the fine-tuning experiment. At the core, Kavier first computes the step runtime and then derives the throughput.
Time per step¶
A four-term sum: the forward pass, the backward pass, the optimizer update, and the gradient synchronization across GPUs.
\(G_a\) is the number of gradient-accumulation micro-steps (default \(1\)).
Forward pass and MFU¶
\(P\) is the number of model parameters used in the forward pass (the active parameters of an MoE model), \(B\) the micro-batch size, \(S\) the sequence length in tokens, \(F\) the GPU's peak FP16 Tensor Core throughput, \(E\) the effective MFU, and \(O_t\) the calibrated training overhead. \(E_b\) is the GPU's nominal MFU factor, \(E_g\) a per-GPU calibration factor, and \(a_1, a_2\) are calibration factors.
Backward pass¶
Optimizer update¶
AdamW, dominated by memory traffic: 20 bytes moved per trainable parameter, each step.
\(B_m\) is the GPU memory bandwidth, \(P_{all}\) the total number of model parameters, \(r = 8\) the LoRA rank, \(d\) the hidden size, \(k = 4\) the number of target modules per layer, and \(L\) the number of transformer layers.
Gradient communication¶
A ring all-reduce: none on one GPU, one ring at intra-node bandwidth on one node, and an intra-node then inter-node (InfiniBand) all-reduce across nodes.
\(G\) is the total number of GPUs, \(N\) the number of nodes, and \(G_n = \max(1, \lfloor G / N \rfloor)\) the GPUs per node. \(W_n\) is the GPU interconnect rate and \(W_i = 200\) Gbit/s the InfiniBand rate. \(D = 4 \times P_t\) is the gradient size in bytes, \(\ell\) the per-hop latency, \(o\) the per-message overhead, and \(c_c\) a calibrated communication scale.
Throughput¶
\(c_m\), \(c_o\), and \(c_i\) are calibration scales per fine-tuning method, per LLM, and per (model, method, GPU, \(G\)) interaction; \(m_g\) is a multi-GPU correction.
Calibration¶
Pure physics-driven modeling is not enough when simulating real-world large-scale systems, so Kavier is
calibrated in two tiers on LLMFineTuningBench:
- Global: one correction factor per GPU model, fine-tuning method, LLM, GPU count, and one for communication, fitted together on the training split (70%) with Powell's method.
- Four-way: for each (LLM, method, GPU type, GPU count) group, the median ratio of measurement to the tier-1 prediction.
On the held-out test split, calibration reduces the MdAPE, averaged over the four dense models, from 12.5% to 5.4% (MSc thesis, E1).
Calibration is keyed on exact catalog names. An uncalibrated name falls back to a neutral 1.0.