About
We learned this by running our own AI in production.
Most firms selling AI have never had to keep one running. We built a multi-tenant AI product of our own, put it in front of paying customers, and have been on the wrong end of its incidents. Everything we now do for clients came out of that.
What we own
A production AI platform, not a portfolio of pilots.
The clearest thing we can tell you about how we build is that we already run something ourselves, at scale, with money attached.
GenAI Ranker is our product. It measures how six AI engines describe a brand against its named competitors: multi-tenant, row-level tenant isolation, orchestration across six model providers, live billing, paying customers, and real incident nights.
Everything we sell came out of operating it. Eval suites exist because a model update broke something for us before it broke something for a client. Cost ceilings exist because we have watched a token bill behave badly. Refusal design exists because a confident wrong answer is worse than no answer, and we found that out on our own product rather than on yours.
It is also why we can run AI search visibility for clients at almost no marginal cost, and why the first scan is free rather than a lead magnet with a price behind it.
Most firms selling AI can show you a prototype. We can open a production system in the meeting and let you click around it.
How we build
The parts that decide whether it survives.
Any competent team can produce a demonstration. These are the six things that determine whether a system is still accurate, affordable and trusted eighteen months later — and they are what we would want to be assessed on.
Nothing goes live without an accuracy baseline
Before a system speaks to a customer or writes to a financial record, it is tested against a labelled set drawn from your own data, with a pass threshold agreed in writing. You get a measured figure every month afterwards, not an assurance.
Regressions are blocked in CI, not discovered by users
Every build ships with an eval suite. When a model changes underneath us — and providers deprecate on their own schedule — we re-run it, see exactly what moved, and migrate deliberately. Without that suite you find out from a complaint.
Refusal is designed before the happy path
A system that says it does not know is worth more than one that guesses convincingly. Retrieval cites its source; anything ungrounded is refused and escalated. We test that behaviour first, because it is the one that protects you.
Every decision is traceable
Input, retrieved context, model and version, guardrails applied, human approval. That record is as much the deliverable as the software, and it is what gets a system through a risk review rather than stuck in one.
Cost is architected, not discovered
We model cost per transaction before writing code — caching, routing, right-sized models, hard ceilings. Model and infrastructure costs are billed at cost against a ceiling you approve. A margin on tokens would corrupt every one of those decisions.
Oversight sized to consequence
Approval gates and escalation paths scaled to what an error actually costs, not to how the demonstration looks. Where being wrong is expensive, a person signs off, and the sign-off is recorded.
Why us
Most AI projects die after the demo.
We ship production AI — including our own
GenAI Ranker is our product: multi-tenant, six model providers orchestrated, row-level tenant isolation, live billing and real incident nights. Most firms selling AI can show you a prototype. We can show you a production system we own, operate and are paged for — and you can watch us open it in the meeting.
We build with agentic AI — fast, and still disciplined
Our own delivery runs on agentic tooling, which is why an agent is live in days rather than a quarter. The speed does not come out of the quality budget: eval suites, regression gates, tracing, guardrails and cost ceilings are still there, because that is what makes it survive contact with real users.
We argue about the business case, not just the architecture
We cost the current process before proposing a system, and we can defend the build to a finance committee as well as to an architecture review. Most AI vendors are fluent in one of those conversations. The one your board will actually hold is usually the other.
We never mark up model or infrastructure costs
Billed at cost against a ceiling you approve. A margin on tokens would corrupt every architectural decision we make on your behalf — model choice, caching, routing, context size, all of it.
We build the demand layer too
A platform nobody arrives at and an agent nobody talks to are worth nothing. We run the search, content and campaigns that feed what we build, instrumented end to end — so the system and the demand for it are designed together.
We keep it running
Monitoring, retraining, eval regressions, model migrations and cost control after launch. Most firms hand over and vanish, right before the model underneath gets deprecated. That gap is the whole reason clients stay with us.
The name
Geeks & Nomads.
It describes the two halves of the job. The geek builds the system: precise, tested, instrumented, cheap to run. The nomad goes where the work is and adapts when the ground moves — which in this field is roughly every quarter.
The mark is two open brackets rotating around a node. Violet is build, orange is run. They never close, because the system is never finished.
Senior people on the work, not just on the pitch.
On every engagement, the people who write the plan are in the delivery, and are reachable when it matters. That is the model, not a favour we extend to large accounts.