The largest language models get the headlines, but many production workloads run better, faster and cheaper on smaller models. Choosing model size deliberately is one of the most effective ways to control generative-AI cost without sacrificing quality.
What counts as “small”?
There’s no official cut-off, but small language models (SLMs) typically have a few billion parameters or fewer, while frontier large language models (LLMs) are far larger. Most providers offer families with several sizes — a fast, inexpensive tier and a most-capable tier.
Where large models win
- Complex, multi-step reasoning and planning — including most agentic workflows.
- Broad, open-ended questions across many domains.
- High-stakes writing and analysis where nuance matters.
- Tasks with little task-specific training data.
Where small models win
- Well-defined, repetitive tasks: classification, extraction, routing, short summaries.
- High-volume workloads where cost per request dominates.
- Low-latency experiences such as autocomplete or real-time assistance.
- On-device, edge or air-gapped deployments where data cannot leave the environment.
- Tasks where fine-tuning on your data can close the quality gap.
Side-by-side
| Small models | Large models | |
|---|---|---|
| Cost per request | Low | Higher |
| Latency | Low | Higher |
| Complex reasoning | Limited | Strong |
| Self-hosting | Practical on modest hardware | Usually via cloud APIs |
| Fine-tuning | Cheap and fast | Costlier |
The best answer is often both
Many production systems route requests: a small model handles classification, extraction and simple questions, and only complex requests go to a large model. Others use a large model to generate training examples that fine-tune a small model for a narrow task. Both approaches can cut costs substantially while keeping quality high.
How to decide
- Build an evaluation set from your real task.
- Test a small, medium and large model on quality, latency and cost.
- Pick the smallest model that meets your quality bar.
- Re-test periodically — small models improve quickly.
Our AI & Machine Learning team helps you benchmark models on your own tasks and design cost-efficient architectures.
Frequently asked questions
What is a small language model?
A small language model is a compact language model — typically a few billion parameters or fewer — that is cheaper and faster to run and can often be self-hosted, at the cost of weaker complex reasoning.
Are small language models good enough for business use?
For well-defined tasks such as classification, extraction and routing, often yes — especially when fine-tuned. Complex reasoning and agentic tasks usually still benefit from larger models.
How do I reduce LLM costs?
Use the smallest model that meets your quality bar, route simple requests to cheaper models, keep prompts and context lean, cache repeated work and monitor usage.