Less Is More: Why I Reach for Simpler ML Models First
Why SVMs, Mask R-CNN, and classical ML models still outperform deep learning in production, and how to choose the right model for accuracy, latency, and total cost of ownership.
There’s a pattern in ML engineering that I see constantly: a team reaches for the most sophisticated model available, spends weeks fine-tuning it, and ends up with something that’s slower, harder to maintain, and barely more accurate than a solution they could have built in a day.
I start simple and only add complexity when the data demands it.
SVMs are not obsolete
Support vector machines have been around since the 1990s. In the age of large language models and diffusion networks, they feel almost quaint. But for a surprising number of classification problems, especially with structured, tabular data, an SVM with a well-chosen kernel will match or beat a deep learning classifier while training in seconds instead of hours.
Consider a binary classification task on a dataset with 50 features and 10,000 samples. A gradient-boosted tree or SVM will excel on these tasks and give you 95%+ accuracy with minimal tuning. A neural network might get you to 96% or higher, but now you need GPU infrastructure, hyperparameter sweeps, and a training pipeline that takes 100x longer to iterate on.
The math isn’t complicated: if the marginal accuracy gain doesn’t justify the infrastructure and maintenance cost, the simpler model wins. For most business problems, less is more.
SVMs also have properties that matter in production. They’re deterministic, they run inference in microseconds, and they have well-understood theoretical guarantees about generalization. They’re trivial to deploy: serialize the model, load it in any language, call predict. There’s no GPU, model server, or CUDA driver to manage.
Mask R-CNN still holds up
Instance segmentation is a good example of where the “latest model” instinct can mislead you. YOLO and SOLO variants have gotten very good: fast inference, solid accuracy, easy to train. For real-time applications where you need 30+ FPS on edge hardware, they’re the right choice.
But when accuracy is the priority, as in medical imaging, industrial inspection, or satellite analysis, Mask R-CNN is still a perfectly viable solution. Its two-stage architecture (region proposal, then classification and mask prediction) gives it an edge on complex scenes with overlapping objects, unusual aspect ratios, or fine-grained boundaries.
I’ve deployed Mask R-CNN pipelines in environments where the cost of a missed defect is orders of magnitude higher than the cost of an extra 200ms of inference time. In that context the model’s thoroughness is the reason to use it. The two-stage approach catches things that one-shot detectors skip.
Tip
The right model depends on the constraints. Real-time tracking in a warehouse? YOLO. Inspecting semiconductor wafers where a missed defect costs thousands? Mask R-CNN’s accuracy premium is worth every millisecond.
The modern model trap
The ML community, myself included, has a bias toward novelty. New architectures get published, benchmarks get topped, and suddenly last year’s state-of-the-art feels outdated. But benchmarks measure performance on benchmark datasets, which aren’t your data, your constraints, or your production environment.
I’ve seen teams spend months implementing vision transformers for problems that a ResNet-50 would have solved in a week. I’ve seen startups deploy GPT-4 for classification tasks where a fine-tuned logistic regression on TF-IDF vectors would have been faster, cheaper, and more predictable.
The pattern is always the same: the team optimizes for model sophistication when they should be optimizing for time-to-production and total cost of ownership. More often than not, the model’s complexity is a distraction from the core problem at hand.
How I choose
My model selection process is deliberately boring:
-
Define the constraints first. Latency budget, accuracy threshold, hardware limitations, data volume, update frequency. These narrow the field before I look at a single architecture.
-
Start with the simplest viable model. For tabular data, that’s usually a gradient-boosted tree or SVM. For image tasks, a pre-trained CNN. For text, TF-IDF plus a linear classifier. Train it in an afternoon, evaluate it earnestly.
-
Only add complexity if the simple model falls short. If the baseline meets the accuracy threshold and fits within the latency budget, I ship it. If it doesn’t, I have a concrete gap to close instead of a vague sense that I should be using something fancier.
-
Account for the full lifecycle. A model that’s 2% more accurate but requires a GPU cluster, a dedicated ML engineer to maintain, and a 6-hour retraining pipeline is often worse than the simpler alternative. I factor in deployment, monitoring, retraining, and debugging costs from day one. 10x the engineering effort is frequently not worth 10x the budget.
Complexity has a tax
Every layer of complexity in an ML system creates added maintenance burden. Deep learning models need GPU infrastructure. Large models need model serving frameworks. Custom architectures need engineers who understand them. Each of these is an ongoing cost that compounds over time, and can shrink your bus factor to one.
A scikit-learn model deployed as a Python function behind a REST API is something any backend engineer can debug, retrain, and redeploy. A PyTorch model running on a Triton inference server with custom CUDA kernels is a different story.
I’m not anti-deep-learning. I use transformers and complex architectures when they’re needed. But in most real-world applications the right answer is simpler than you’d expect, and the discipline to ship the simpler solution is what separates production ML from research ML.
The cheapest model that meets the bar
The best model for your problem is the one that meets your accuracy requirements with the lowest total cost of ownership. Sometimes that’s a frontier LLM. Often, it’s an SVM trained on a 10-year-old ThinkPad in 30 seconds. Engineering expertise is knowing which is which, and shipping the boring solution when it’s the right one.
This model-selection discipline is the backbone of how I approach computer vision and AI integration work: start simple, measure, and only add complexity the data demands.
Content may be edited with the help of AI. All content has been reviewed by a human before being published.
Have a project in mind?
I design and ship custom software — from early concept to production. Tell me what you're building and we'll figure it out together.