Llama 3.1 8B Instruct
NVIDIA A100 (40 GB), recorded Jul 12, 2026
- Time to first token
- 71.12 msp50, batched serving (FP16, batch 24)
- Output throughput
- 391.76 tok/smeasured, batched serving (FP16, batch 24)
- Per 1M output tokens
- $1.489derived from recorded GPU rate and measured saturation, batched serving config
- Quality vs FP16 baseline
- 0.973reported candidate, precision unverified
Historical record llama31-8b-a100-2026-07-12, incomplete provenance