Knowledge and document assistants
Answering questions over policies, contracts, manuals or case history, with citations back to the source so an answer can be checked rather than trusted.
RAG · Citations · Access controlCrux builds AI applications that hold up under real inputs — grounded in your data, measured against an evaluation set, and constrained so a wrong answer is caught rather than shipped.
Most things described as AI applications are a form with a model call behind it. That works in a demo because the demo uses inputs the builder chose. Production supplies inputs nobody anticipated, at volumes nobody tested, from users who will paste anything into a text box.
An AI application is the engineering around the model call: retrieval that gives it the right context, validation of what comes back, evaluation that catches regressions before users do, and a defined behaviour for the cases where the model should decline rather than answer.
That surrounding system is most of the work and nearly all of the risk. The model is a component you can swap; the application is what makes it dependable enough to put in front of a customer or a regulator.
Getting this wrong is the most expensive mistake available, and it is made before any code exists. The honest default is to buy generic capability and build only where your data or process creates the advantage.
| Approach | Choose it when | Real cost | Verdict |
|---|---|---|---|
| Buy a product | The capability is generic — transcription, translation, standard document extraction — and a vendor already does it well. | Licence, plus lock-in and limited control over behaviour. | Default choice |
| Prompt + retrieval | The value is in your documents or data, and the task is answering, summarising or extracting from them. | Engineering time, retrieval quality work, ongoing token cost. | Start here |
| Fine-tune | Prompting plus retrieval has been measured and falls demonstrably short, usually on format, tone or a narrow domain vocabulary. | A model you now own, evaluate, version and re-do when the base model moves. | Prove it first |
| Train from scratch | Almost never, outside genuinely novel domains with proprietary data at real scale. | Capital expenditure and a specialist team most organisations do not have. | Rarely justified |
Every AI application we ship has these, in some form. Their absence is what distinguishes a pilot from a system.
Finding the right context is where most quality is won or lost. Chunking strategy, embedding choice, hybrid keyword-and-vector search, and reranking matter more than the model, and considerably more in Arabic, where morphology and normalisation affect retrieval quality directly.
Provider and size chosen per task rather than per project. Small models handle classification and extraction at a fraction of the cost; frontier models earn their price only on reasoning-heavy steps. Routing between them is a design decision with a monthly invoice attached.
Structured output validated against a schema before it reaches anything downstream. A model returning malformed JSON should fail loudly, not silently corrupt a record.
Defined behaviour for the edges: refusing when retrieval returns nothing relevant, refusing outside scope, filtering sensitive content, and rate-limiting per user. Refusal is a feature, and one that has to be designed.
A dataset of representative inputs with agreed correct outputs, scored on every prompt or model change. Without it there is no way to know whether last week's improvement broke something else.
Tracing every call with its inputs, retrieved context, cost and latency. When a user reports a bad answer three weeks later, this is the difference between a fix and a shrug.
Most enterprise requests reduce to one of these, or a combination. Naming the pattern early makes scope, cost and risk far easier to estimate.
Answering questions over policies, contracts, manuals or case history, with citations back to the source so an answer can be checked rather than trusted.
RAG · Citations · Access controlTurning invoices, forms, reports and correspondence into structured records, with confidence thresholds routing uncertain extractions to a person.
Extraction · Schema validationDemand forecasting, risk scoring, churn and propensity, delivered as services other systems call rather than dashboards people read.
Forecasting · Scoring APIsTriage at volume: tickets, applications, complaints and correspondence sorted and routed, with the ambiguous minority escalated.
Triage · Routing · EscalationSearch, summarisation, sentiment and conversational interfaces that work in the Arabic people actually write, including Gulf dialect in user-generated text.
Arabic NLP · Dialect · SearchMulti-step execution across systems where each step depends on the last. Different discipline and different risk profile — see our agentic AI work.
Multi-step · Approval gatesAsk a vendor how they will prove their AI application is correct. If the answer is a demo, the project has no quality control. Software has tests; AI applications need evaluation sets, and building one is the first engineering task rather than the last.
This is also the honest answer to the hallucination question. It is not solved by a better prompt. It is managed by grounding answers in retrieved sources, validating structure, designing refusal, routing consequential actions to human approval, and measuring continuously.
AI applications have a marginal cost per use, which is unfamiliar to organisations accustomed to licensing software once. A feature that is delightful at a hundred users a day can be indefensible at a hundred thousand.
Both budgets belong in the design conversation. What is the acceptable cost per interaction, and the acceptable wait? Those two numbers drive model selection, routing, caching and how much retrieval you can afford per query — and they are far cheaper to answer at design time than to retrofit after launch.
Which pattern this is, whether the data supports it, and what "correct" means for this task. Ends with a build/buy recommendation, including the recommendation not to build.
Representative inputs and agreed correct outputs, built with the people who own the process. Everything afterwards is measured against this.
Retrieval, model layer, validation, guardrails and integration. Scored against the evaluation set continuously rather than at the end.
Adversarial testing, latency and cost tuning, PDPL review of data flows including prompts sent to third-party providers, and human review workflows.
Deployment with tracing and monitoring in place, operator documentation, and handover to the team who will run it.
Full enterprise applications run 3 to 6 months. The constraint is almost always data readiness and integration count — rarely model capability.
Buy when the capability is generic and a vendor already does it well. Build when the value comes from your data or your process. Fine-tune only when prompting plus retrieval has been tried and measurably falls short — fine-tuning creates a model you now own, must evaluate, and must redo when the base model moves on.
Retrieval-augmented generation fetches relevant documents at query time and gives them to the model as context. It suits questions answered by a body of knowledge that changes: policies, contracts, product documentation, case history. Usually cheaper, more current and more auditable than fine-tuning, because you can show which source produced an answer.
Through an evaluation set built before the application: representative inputs with agreed correct outputs, scored automatically on every change. Without it, quality is whatever the last person to try it thought, and every model change is an unmeasured risk.
By constraining what the model is asked to do. Grounding answers in retrieved sources with citations, refusing rather than guessing when retrieval returns nothing relevant, validating structured output against a schema, and routing consequential actions to human approval. Managed by design, not eliminated by a better prompt.
Yes, with deliberate handling. Modern Standard Arabic is well supported and Gulf dialect is workable. Retrieval is the usual weak point, since Arabic morphology and normalisation affect chunking and embedding more than in English. Arabic evaluation sets are built alongside English ones rather than assumed equivalent.
A first production release typically takes 12 weeks including the evaluation harness. Full enterprise applications run 3 to 6 months, constrained by data readiness and integration count rather than model capability.
They can be, if designed that way. Lawful basis, data residency, retention and audit logging affect architecture — and prompts sent to third-party model providers are themselves a data transfer that has to be assessed. Where policy requires it, models run inside your environment.
Multi-step workflows where an agent acts across systems, with authority limits and approval gates.
Explore → ConnectConnecting AI capability to the ERP, CRM and data systems where decisions are actually made.
Explore → DecideWorking out which problems are worth solving with AI before anyone writes a line of code.
Explore →A two-week scoping engagement establishes which pattern your problem fits, whether your data supports it, what "correct" means, and what it will cost per interaction at your volume — including the recommendation to buy or to do nothing.