1. Author
    Célio Pires
  2. Published
    2026
  3. Category
    AI & Design

Shipping AI, not just using it

Using AI in your own workflow makes you faster. Shipping AI-powered systems makes a whole team better. The difference is where the judgment lives.

Two different jobs

Most designers I talk to use AI every day. They write prompts, generate variations, summarise research, draft copy. That's useful. It's also personal. When they close the laptop, the value leaves with them.

Shipping AI is a different job. You build something other people run. An agent that reviews every change the same way. A research engine that maps context before anyone designs. A connector that lets a customer get real work done from a chat window.

The first job is about speed. The second is about systems. And systems are what I've been doing my whole career.

Encode judgment, not prompts

A prompt is a request. An agent is a standard.

OnEnsinova I built a set of specialist agents: a UX reviewer, a UI-governance agent that enforces the design system, a responsive tester that checks every change at four screen widths, and a "dead-ends" agent that hunts for screens with no way forward.

None of them are clever on their own. They're useful because each one carries a slice of judgment I used to apply by hand in design reviews. Now it runs every time, whether I'm in the room or not.

Write down the rules

The agents only became useful when I wrote down what I'd learned the hard way.

• every incident becomes a rule in the agents' shared memory
• every decision gets a written spec with the reasoning behind it
• every request gets acceptance criteria that can actually be graded

This is design systems work. Tokens, components and documentation were always a way to encode judgment so other people didn't have to rediscover it. Agents are the same idea, one layer up.

Where trust lives

When I shipped the Ensinova MCP connector, school owners could connect their centre to Claude or ChatGPT and ask "who hasn't paid September?" against live data. The hard part wasn't the model. It was deciding where trust lives.

A few architectural decisions shaped everything:

•The model is an untrusted client. Who you are and which school you belong to come from the login token, never from what the model sends. A prompt can't move a session into another school.
•Invisible, not denied. Tools are registered per role. A teacher's assistant never even sees financial tools, so it can't reason its way into them.
•Wide reads, narrow writes. Reads use a read-only database role. Writes go through the product's own API, with the same permission checks as the UI. One source of truth, not two.
•Confirmation as interface. Every write starts as a preview a human confirms. A retry never creates a student twice.

These are design decisions as much as engineering ones. You can't make them well unless you understand the parts underneath: tokens, permissions, tool registration, what the model can and can't see.

Discovery as a system

Discovery scales when it's a system, not a ritual.

Before anything gets designed, a researcher agent maps the request against the codebase, the data model and earlier decisions. A clarity gate sends vague requests back with questions instead of guessing. What comes out is a spec with testable criteria, not a ticket.

The same thinking runs the help centre: an engine that captures every documented flow in light and dark mode and regenerates it whenever the UI changes.

The best tooling is invisible

The best design system is the one people use without thinking about it. AI tooling works the same way.

A designer shouldn't need to know how the review agent was built to benefit from it. They open a change, the standards are already applied, and the reasoning is written down if they want it.

If people have to learn your system before they can use it, you've built a tool for yourself.

Guardrails, not hope

Autonomy without limits is just risk.

My pipeline can never push code or mark work as done. That's blocked in the instructions and in the tool permissions. Every task ends in "In Review", for a human. A pre-commit hook blocks unsafe data access and leaked secrets.

I don't trust the agent to behave. I design the system so it can't misbehave in the ways that matter.

What I learned

AI doesn't replace good judgment. It amplifies the judgment you already have, and only if you write it down.

The skills that made me good at design systems turned out to be the skills that matter most here: encoding decisions, designing for people who will never meet you, and caring about adoption as much as the thing itself.

Using AI makes you faster. Shipping it makes everyone around you better.

Latest Articles

Insights and learnings from building design systems across different products and teams.View all articles