Ethics in AI: The Right Way Is the Hard Way
Why I avoid diffusion models, train exclusively on authorized data, and believe AI works best when applied surgically, not as a blunt instrument.
People are right to be skeptical about AI, in the concrete, “my work was used to train a model without my consent” sense. The concerns around AI ethics right now are not hypothetical. They’re grounded in real harm, active lawsuits, and an erosion of trust between technology companies and the people their products are supposed to serve.
I understand these concerns because, even as someone who builds AI systems for a living, I share them.
The diffusion model problem
The most visible AI ethics issues today center on generative media models: image diffusion models, video generators, and music generators. These systems are trained on massive datasets scraped from the internet, often without the knowledge or consent of the people who created that data. Artists find their styles replicated, or their art generated outright. Photographers find their images used, without consent, to train systems that compete directly with them and sometimes reproduce their work outright.
The standard industry response is some variant of “fair use” or “transformative work.” I don’t find that argument particularly compelling. The fact that a model doesn’t store exact copies of its training data doesn’t mean the people who created that data aren’t essential to the model’s capabilities. Laundering influence through billions of parameters doesn’t make it ethical.
This is why I don’t build or deploy diffusion models, and why I avoid generative media models entirely. My work focuses on discriminative models: classifiers, detectors, and segmentation networks. Models that analyze and categorize rather than create. The distinction matters because it changes the relationship between the model and its training data entirely.
Training data matters
When I train a model, I use exclusively authorized or otherwise publicly available and permissively licensed data. It’s a constraint I design around from the start.
This means I work with datasets that have clear provenance. Data that was collected with informed consent, released under permissive licenses, or generated specifically for the task at hand. When I work with client data, the terms are explicit: what the data is used for, how long it’s retained, and what happens to it when the engagement ends. If clients don’t want their data used in AI models, I don’t use it.
Note
Using authorized data is also a practical choice. Models trained on clean, well-labeled, domain-specific data consistently outperform models trained on massive but noisy internet scrapes for the kinds of focused problems I solve.
This approach is more constrained than scraping the entire internet, and the constraints lead to better engineering. When you can’t brute-force your way to performance with more data, you’re forced to be smarter about feature engineering, model selection, and problem framing. The result is leaner models that perform better on the actual task, which is different from a benchmark designed to reward scale.
Applied surgically
There’s a prevailing approach to AI right now that treats it like a universal solvent. Just pour a large enough model onto any problem and it’ll dissolve. Need customer support? Chatbot. Need image editing? Diffusion model. Need to summarize a document? Another billion-parameter model. The solution is always more parameters, more compute, more data.
I think this is backwards. AI works best when it’s applied surgically: the right model, at the right scale, for a specific problem with well-defined constraints. An SVM that classifies defects on a production line. A Mask R-CNN that segments regions of interest in medical imaging. A random forest that flags anomalous transactions. These applications aren’t glamorous, and they work.
The “AI as a bludgeon” approach creates problems beyond ethics: dependency on massive compute infrastructure, systems that are opaque and difficult to debug, and solutions overengineered for the task at hand, with the maintenance burden that implies. It also concentrates power in the handful of companies that can afford to train and serve models at that scale.
And frankly, it’s just annoying when your favorite app gets an AI assistant forced into it for no good reason at all.
AI as a force for good
None of this means AI is inherently harmful. The technology itself is a powerful tool, and tools can be wielded responsibly or recklessly.
AI is a force for good when it makes medical diagnoses faster and more accessible. When it helps farmers detect crop disease before it spreads. When it makes industrial processes safer by catching defects that human inspectors miss. When it helps small teams punch above their weight by automating the tedious parts of their work so they can focus on the parts that actually require human judgment.
But these applications share a common trait: they’re specific, targeted, and deployed with a clear understanding of what the model does and doesn’t do. They treat AI as a precision instrument.
Where I draw the line
My position is simple. I build AI systems that:
- Don’t generate content that could displace creative work. I build models that analyze, classify, and detect, not models that produce images, video, or music.
- Train exclusively on authorized data. If I can’t verify the provenance and licensing of a dataset, I don’t use it.
- Solve specific, well-defined problems. I don’t build general-purpose AI systems. I build tools that do one thing well, with clear boundaries and predictable behavior.
- Remain explainable. When a model makes a decision, I can tell you why. This matters for effective model development, and for establishing trust and accountability.
These are engineering constraints that shape every project I take on. They make some things harder and some things impossible, and I’m fine with that, because the things they make impossible are things I don’t want to build anyway.
The trust problem
The AI industry has a trust problem, and it earned it. When companies scrape the internet without consent, deploy opaque systems that affect people’s livelihoods, and dismiss legitimate concerns as technophobia, they shouldn’t be surprised when the public pushes back.
I’m not here to rehabilitate AI’s reputation. I’m here to build things that work, for people who need them, using methods I can defend. If the industry eventually lands on meaningful standards for ethical AI development, great. In the meantime, I’m not waiting for permission to build AI systems the right way.
These principles shape every computer vision and AI system that ships under the Arclight name.
Content may be edited with the help of AI. All content has been reviewed by a human before being published.
Have a project in mind?
I design and ship custom software — from early concept to production. Tell me what you're building and we'll figure it out together.