The interesting question is what it is allowed to do
Generating text is solved and nearly free. The difficulty is connecting a model to your actual data and letting it act on your actual systems without anybody getting hurt — and that is an engineering problem about permissions, verification and rollback rather than a prompting problem.
YOU ARE PROBABLY HERE BECAUSE
- you are doing something tedious that a machine could read
- customers ask an assistant before they visit your site
- somebody quoted you for AI and you cannot tell if it is real
What I can do for you
The first one exists because most AI projects fail on the question of what should be automated, not on the model.
Feasibility and cost model
What you want, whether a language model is the right tool for it, and what it would cost per request at your volume. Includes the cases where the honest answer is that a database query or a rules engine does the job better, cheaper and deterministically.
AI feature build
An assistant, a search that understands questions, extraction from documents, classification, summarisation, drafting. Built into the application you already have, grounded in your own data, with the cost and the failure behaviour designed rather than discovered.
Agentic systems
Software that takes actions rather than only producing text — with tools, permissions, approval gates for anything irreversible, and a full audit of what it did and why. Also exposing your own systems to agents through MCP, so other people's assistants can use them.
Making your site usable by agents
Starting with a crawler reachability audit, which is a day and frequently finds the whole problem. Then structured data, a knowledge catalog so nothing is maintained twice, and an MCP endpoint where there is genuinely something for an assistant to do.
Keeping it working
Models change, providers deprecate, prompts drift, and costs move. Reserved time for evaluation against a fixed test set, cost monitoring, and updating when something upstream changes underneath you.
Forty-five minutes, no charge.
Bring the task you are hoping to automate and I will tell you whether it needs a model at all.
01
How much autonomy to give it
Almost every decision worth making happens here. The further right you go, the more useful and the more expensive a mistake becomes — and the line between suggesting and acting is where most of the engineering lives.
02
What has become worth building
Features that were research projects five years ago are now a few days of work. Not because the ideas changed, but because the hard part is now a service call.
A customer asking "does this fit a 2019 model" gets an answer from your own catalogue and documentation, rather than a results page containing the word "2019". Grounded in your data, with citations, so it cannot invent products you do not sell.
Purchase orders, invoices, specifications and supplier price lists turned into structured data. Previously either manual typing or a brittle template parser that broke whenever a supplier changed their layout.
Incoming enquiries, tickets or orders sorted, tagged and sent to the right place. Unglamorous, easy to evaluate, and usually the highest return of anything on this list.
Product descriptions from specifications, replies from a support history, translations, summaries of long threads. A person still approves, which is what makes it safe and what makes it fast.
Staff asking questions of documentation, policies, order history or stock, in plain language. Internal ones are considerably easier to justify than customer-facing ones, because the cost of an occasional wrong answer is much lower.
Reformatting, categorising, tagging or translating a large catalogue — work that was previously an outsourcing project and is now a weekend of compute with a human sampling the output.
Exposing your services, availability or catalogue so an assistant acting for a customer can query and use them. Increasingly your buyer asks a model before visiting your site. See how this site does it →
A sequence rather than a single answer — read the enquiry, look up the account, check stock, draft the reply, queue it for approval. Each step is a real function with real permissions, and the model decides the order rather than doing the work itself.
03
What this does to a project's cost
Two separate things get called cost saving and they behave differently. It is worth being precise about which one you are buying.
These tools have genuinely changed the pace of development, most visibly on the boring middle of a project — scaffolding, tests, migrations, documentation, unfamiliar APIs. It shows up in my estimates rather than in a line item, because it is a way of working rather than something I sell you.
Understanding a fifteen-year-old codebase, deciding what should be built, and the last twenty percent of anything touching money still take exactly as long as they did. Anyone quoting you dramatic savings on those has not done the work.
The larger saving is not building the same thing faster — it is that things which would have needed a specialist team and a research budget are now a normal feature. Semantic search, document understanding and language handling are the obvious examples.
Unlike ordinary code, these features cost money every time they run. That has to be modelled at your real volume before you commit, and designed for — caching, smaller models for simpler steps, and not calling a model where a query would do.
Providers deprecate models, behaviour shifts between versions, and a prompt that worked in March may not in September. Budget for evaluation and upkeep, or expect the feature to quietly decay.
04
Platform and language agnostic
These are HTTP APIs. The interesting work is in your application, not in which language calls the model — and the same patterns apply wherever it plugs in.
WordPress and WooCommerce are where I am deepest, but a PHP, Laravel, Node or Python application is the same job. Most of this work is integration into something that already exists rather than a new build.
Built behind an abstraction so the provider can change without rewriting your feature, because prices and capabilities move quickly and the best option in six months will not be today's.
Hosted APIs, a provider with the right data-handling terms, or a model running on your own infrastructure. Regulated and sensitive work sometimes rules out the easy option entirely, and that decision comes first, not last.
Most business cases are answered by giving the model your documents at request time rather than training a custom model. Cheaper, updatable the moment your content changes, and it can cite where an answer came from.
Where your systems need to be usable by assistants, that is better done with the open protocols emerging for it than with a bespoke integration per vendor.
05
Making your business usable by other people's agents
Increasingly your customer asks an assistant before they visit your site. If it cannot read what you offer, cannot check your availability and cannot start the transaction, you are not in the answer — and no amount of design fixes that, because nothing with eyes is looking at your page.
This is a stack of layers, and they are not equally worth having. I would rather tell you which ones are cheap theatre than sell you all of them.
Whether AI crawlers and assistant fetchers can reach you at all. Managed hosts and edge providers increasingly block them by default — and that happens before robots.txt is read, so a perfect robots file changes nothing. I check from outside with the actual user agents rather than trusting a dashboard.
A page that returns real content without executing JavaScript, with correct headings, tables and lists. Assistant fetchers behave much like a simple crawler. A site that renders entirely in the browser is invisible to them, however good it looks.
Your services, prices, availability, FAQs and reviews expressed as machine-readable facts. Generated from one source so the markup cannot drift away from the page it describes — which is what happens when somebody maintains JSON-LD by hand.
A single store of what your business knows — services, projects, policies, availability, canonical answers — with a disclosure level on every entry so confidential work can be counted without being named. Every output renders from it: the page a person reads, the JSON-LD, the description files, the tools an agent calls. Nothing is ever typed twice, so nothing can disagree.
Description files at well-known URLs. They cost almost nothing to generate from the catalog, and the evidence for them is weak — adoption sits at roughly two percent of sites and a fraction of a percent of AI bot requests target them. Worth publishing as an operating manual for an agent. Not worth paying anyone much for.
The genuinely different layer. An endpoint exposing real functions an assistant can call — search your catalogue, check availability, look up an order, start a booking. Built on the WordPress Abilities API where the site is WordPress, or as a standalone service where it is not. Read-only first, writes behind the same approval design as everything else here.
Where MCP lets an assistant use your tools, A2A lets another organisation's agent delegate a task to yours — a supplier's system asking yours to quote, with a long-running task that can pause and ask a clarifying question. Worth it across a company boundary; overkill for most single sites, and I will say which yours is.
An assistant on your site answering questions from the catalog in your own voice, and able to actually do things — check stock, start a booking — because the tools already exist by this point. Built last deliberately: an assistant with nothing solid underneath it is a chat box that makes things up.
Logging which AI crawlers and fetchers reached you, what they took, and which agent tools were called. Without it there is no way to know whether any of this worked — and it is the only honest basis for deciding what to do next.
What goes wrong
Most of these are design failures rather than model failures, which is the good news — they are preventable.
Confident and wrong
A model states something untrue in the same tone as something true. Mitigated by grounding answers in your data and citing sources, not by asking it to be careful.
Instructions hidden in content
Where a model reads user-supplied text — an email, a document, a web page — that text can contain instructions aimed at it. Anything with real permissions has to be designed assuming its input is hostile.
No way to tell if it is working
Deployed with no test set, so nobody notices when quality degrades. Traditional tests do not apply, which means building an evaluation set of real cases with known good answers before launch, not after.
The bill nobody modelled
A feature that costs pennies in testing and a great deal at real volume, because per-request cost was never worked out against actual traffic.
Too much autonomy too early
Given permission to act before anyone has established it acts correctly. The approval gate is not a limitation to remove later — it is what makes the first version shippable.
Invisible at the edge
Paying for AI visibility work while the CDN blocks every AI crawler by default. It fails silently — no error, no report, you simply never appear. Ten minutes to check and almost nobody has.
A model where a query would do
Using a language model to do something a database, a regular expression or a rules engine does perfectly, deterministically and free. Common, expensive, and the first thing I check.
What this does not include
Training your own model
Almost nobody needs this, and the people who genuinely do are not reading a WordPress developer's website. Retrieval over your own data solves the business cases.
Replacing your staff
Not what this is good at, and not a project I want. The reliable wins are removing tedious work from people who have better things to do.
Autonomous anything touching money
Payments, refunds, pricing changes and irreversible customer-facing actions go behind a human. I will build the approval step and I will not build past it.
Promising a specific accuracy
I can build an evaluation set and show you how it performs on your real cases. I cannot promise a percentage before seeing your data, and neither can anybody else.
Questions I get asked
Do we actually need AI for this?
Frequently not, and I will say so. If the task has a right answer that can be looked up or calculated, a query or a rules engine is faster, cheaper and correct every time. Language models earn their cost where the input is messy, unstructured or written by a human — which is a real category, just narrower than the current enthusiasm suggests.
What will it cost to run?
That is part of the feasibility work, because it is genuinely different from ordinary software. Cost depends on how much text goes in and out, how often, and which model. It can usually be reduced substantially by caching, by using smaller models for simpler steps, and by not calling a model where a query would do. You get a monthly figure at your real volume before committing.
Can it work with our private data without that data leaving?
Depends how strict the requirement is. Providers offer terms covering data handling and retention that satisfy many businesses. Where they do not, a model can run on infrastructure you control, at higher cost and lower capability. This decision shapes everything else, so it is settled first.
How do we know it is giving good answers?
By building an evaluation set — real questions from your business with known good answers — and running it whenever anything changes. Without that you are relying on someone noticing, and quality tends to degrade quietly. This is the part most implementations skip and the reason they disappoint after a few months.
Should we build this into our existing system or start fresh?
Into what you have, almost always. These features are most valuable when they can see your real data — your orders, your documents, your customers — and that argues for building where the data already lives rather than beside it.
You said this is newer work for you. Why hire you for it?
Because the failure modes are engineering failure modes, not novel ones. Deciding what a system is permitted to do, designing for hostile input, building in an approval step before an irreversible action, and modelling running cost before committing — that is the same discipline as building a payment system, and I have sixteen years of that. What I will not do is claim a longer track record with these specific tools than I have.
Bring the task, not the technology
Describe the tedious thing you would like to stop doing. I will tell you whether it needs a model, what it would cost to run, and when the answer is that it does not.