MENTARA
Software Services

The hard part was never the model.

Most enterprise AI stalls between a convincing demo and a system anyone will depend on. The gap is evaluation, data access, failure handling and ownership.

The decision in front of you

Almost every organisation now has AI pilots that impressed a steering committee and never reached production. The reason is consistent, and it is not model capability: a demo has to be convincing once, and a production system has to be acceptable every time, including on the inputs nobody thought to try.

Closing that distance is ordinary engineering work — evaluation harnesses, retrieval quality, access control, latency and cost budgets, failure behaviour, human escalation, and someone accountable when it is wrong. It is unglamorous, and it is the entire difference between a pilot and a system.

The failure mode

A demo is a system that has only been tried by people who want it to work.

Pilots are evaluated on curated inputs by people who understand the intent and unconsciously phrase things well. Production is the opposite: adversarial in the ordinary sense, full of edge cases, abbreviations, missing context and users who will not rephrase to help. The accuracy figure from the pilot does not survive that transition, and usually nobody measured it in a way that would have predicted this.

The second failure is retrieval, not generation. Most enterprise use cases are grounded in the organisation's own content, and the quality ceiling is set by whether the right passage can be found — which is a data engineering, permissions and content quality problem. Teams tune prompts for weeks when the actual defect is that the source document is out of date, badly structured, or that half the corpus is invisible to the user asking.

The third is that nobody defined what happens when the system is wrong. Not the error rate — the consequence. Who notices, how, what the user sees, whether the output can be traced to its sources, whether a human can intervene before an action is taken. A system that is right 95% of the time and has no answer for the other 5% is not deployable in a process that matters, and the 5% is where all the real design work lives.

Capability

What MENTARA does.

01AI strategy and readinessWhich use cases are genuinely viable given your data, risk appetite and operating constraints — and, more usefully, which are not, before budget is committed to them.
02Retrieval and knowledge systemsThe unglamorous foundation: content quality, chunking, indexing, permission-aware retrieval and evaluation of whether the right source is actually being found. This is where most quality problems are solved.
03Generative AI applicationsProduction applications with evaluation harnesses, cost and latency budgets, traceable outputs and defined failure behaviour — built to be operated rather than demonstrated.
04Agent-enabled workflowsSystems that take actions, with the scoping, permissioning, human checkpoints and audit trail that taking actions requires. Where a deterministic workflow is correct, we build that instead and say so.
05Responsible AI governanceModel and use-case inventory, evaluation standards, human-oversight design, and the disclosure and record-keeping that emerging regimes increasingly require.
06AI operationsMonitoring for quality drift rather than only uptime, cost control, prompt and model version management, and the feedback loop that turns production failures into evaluation cases.
Approach

How the work runs.

  1. 01 Define the acceptance bar first

    Before building: what accuracy, latency and cost make this worth deploying, how it will be measured, and what the system does when it falls short. If that cannot be answered, the use case is not ready.

  2. 02 Build the evaluation before the feature

    A representative evaluation set drawn from real inputs, including the awkward ones. Without it, every subsequent change is a guess about whether things improved.

  3. 03 Fix retrieval and data access

    Content quality, permissions and retrieval accuracy, measured independently of generation. Most quality gains come from here, and they are the ones that persist across model changes.

  4. 04 Ship narrow, with a human in the loop

    One workflow, one user group, oversight in place, instrumented for the failures you have not imagined yet. Widen only on measured performance rather than on enthusiasm.

Starting points

Where engagements usually begin.

01Use-case triageA portfolio of proposed AI initiatives assessed for feasibility, data readiness, risk and value — with a clear recommendation on what to stop.
02Pilot-to-production assessmentWhy a specific promising pilot has not shipped, and what it would concretely take: evaluation, retrieval quality, controls, operational ownership.
03Retrieval quality reviewIndependent measurement of whether your system finds the right source material, and what is capping quality — usually not the model.
04AI governance foundationInventory, evaluation standards, oversight design and disclosure practice, sized to your regulatory exposure rather than to a framework.
Questions

What buyers ask.

Will you tell us if AI is the wrong tool?

Frequently, and this is the most valuable thing we do in this practice. A large share of proposed AI use cases are better served by a rule, a form, a search improvement, a report or a fixed workflow — cheaper to build, cheaper to run and far easier to assure.

Agentic architectures in particular are often applied to problems where a deterministic workflow is correct, which trades reliability for flexibility nobody needed.

Which models and platforms do you work with?

Whichever fits the constraint — the choice is usually driven by data residency, procurement, latency and cost rather than by benchmark performance, and it changes often enough that architectural lock-in to one provider is a design risk in itself.

We hold no vendor partnership or commission arrangement, so the recommendation carries no commercial interest of ours.

How do you handle our data?

Client code, documents and data are never submitted to services that train on submitted content. Where AI tooling is used in delivery, it is disclosed, the tooling is named, and your policy governs. If your policy prohibits it, we work without it.

For systems we build, data residency, retention and access are engagement decisions made explicitly in scoping rather than inherited from a default.

Can you take over a pilot someone else built?

Yes, and it is a common engagement. The first deliverable is usually an honest assessment of whether the existing work should be extended or restarted — and we have recommended restarting when the evaluation and data foundations were absent, which is not the answer anyone wants to hear.

07

Where MENTARA fits best.

Scope

We build for the production acceptance bar rather than for a board demonstration, which takes longer and is less impressive in month one — worth knowing if a quick showcase is what you need.

Novel model development and foundation model training sit outside our scope — that is a research organisation rather than a delivery firm.

Bring the pilot that has not shipped.

Share the business context, constraints and expected outcome. MENTARA will identify the relevant accountable route.

One partner. One plan. Measurable outcomes.