AI engineering · Saudi Arabia

The demo always works. Production is a different discipline.

Crux builds AI applications that hold up under real inputs — grounded in your data, measured against an evaluation set, and constrained so a wrong answer is caught rather than shipped.

Enterprise AI application architecture showing retrieval, model inference, evaluation and guardrail layers
12 weeks To a first production release, evaluation harness included
Grounded Answers cite sources, and refuse when retrieval finds nothing
Measured An evaluation set exists before the application does
PDPL-aware Residency, retention and third-party prompt transfer assessed
01 — The distinction

Calling an API is not an AI application

Most things described as AI applications are a form with a model call behind it. That works in a demo because the demo uses inputs the builder chose. Production supplies inputs nobody anticipated, at volumes nobody tested, from users who will paste anything into a text box.

An AI application is the engineering around the model call: retrieval that gives it the right context, validation of what comes back, evaluation that catches regressions before users do, and a defined behaviour for the cases where the model should decline rather than answer.

That surrounding system is most of the work and nearly all of the risk. The model is a component you can swap; the application is what makes it dependable enough to put in front of a customer or a regulator.

02 — The first decision

Build, buy or fine-tune

Getting this wrong is the most expensive mistake available, and it is made before any code exists. The honest default is to buy generic capability and build only where your data or process creates the advantage.

Approach Choose it when Real cost Verdict
Buy a productThe capability is generic — transcription, translation, standard document extraction — and a vendor already does it well.Licence, plus lock-in and limited control over behaviour.Default choice
Prompt + retrievalThe value is in your documents or data, and the task is answering, summarising or extracting from them.Engineering time, retrieval quality work, ongoing token cost.Start here
Fine-tunePrompting plus retrieval has been measured and falls demonstrably short, usually on format, tone or a narrow domain vocabulary.A model you now own, evaluate, version and re-do when the base model moves.Prove it first
Train from scratchAlmost never, outside genuinely novel domains with proprietary data at real scale.Capital expenditure and a specialist team most organisations do not have.Rarely justified
03 — Architecture

Six layers around the model

Every AI application we ship has these, in some form. Their absence is what distinguishes a pilot from a system.

  1. Retrieval

    Finding the right context is where most quality is won or lost. Chunking strategy, embedding choice, hybrid keyword-and-vector search, and reranking matter more than the model, and considerably more in Arabic, where morphology and normalisation affect retrieval quality directly.

  2. Model layer

    Provider and size chosen per task rather than per project. Small models handle classification and extraction at a fraction of the cost; frontier models earn their price only on reasoning-heavy steps. Routing between them is a design decision with a monthly invoice attached.

  3. Output validation

    Structured output validated against a schema before it reaches anything downstream. A model returning malformed JSON should fail loudly, not silently corrupt a record.

  4. Guardrails

    Defined behaviour for the edges: refusing when retrieval returns nothing relevant, refusing outside scope, filtering sensitive content, and rate-limiting per user. Refusal is a feature, and one that has to be designed.

  5. Evaluation

    A dataset of representative inputs with agreed correct outputs, scored on every prompt or model change. Without it there is no way to know whether last week's improvement broke something else.

  6. Observability

    Tracing every call with its inputs, retrieved context, cost and latency. When a user reports a bad answer three weeks later, this is the difference between a fix and a shrug.

04 — What we build

Six application patterns that reach production

Most enterprise requests reduce to one of these, or a combination. Naming the pattern early makes scope, cost and risk far easier to estimate.

Knowledge and document assistants

Answering questions over policies, contracts, manuals or case history, with citations back to the source so an answer can be checked rather than trusted.

RAG · Citations · Access control

Document and data extraction

Turning invoices, forms, reports and correspondence into structured records, with confidence thresholds routing uncertain extractions to a person.

Extraction · Schema validation

Prediction and scoring services

Demand forecasting, risk scoring, churn and propensity, delivered as services other systems call rather than dashboards people read.

Forecasting · Scoring APIs

Classification and routing

Triage at volume: tickets, applications, complaints and correspondence sorted and routed, with the ambiguous minority escalated.

Triage · Routing · Escalation

Arabic language features

Search, summarisation, sentiment and conversational interfaces that work in the Arabic people actually write, including Gulf dialect in user-generated text.

Arabic NLP · Dialect · Search

Agentic workflows

Multi-step execution across systems where each step depends on the last. Different discipline and different risk profile — see our agentic AI work.

Multi-step · Approval gates
05 — Evaluation

How we know it works

Ask a vendor how they will prove their AI application is correct. If the answer is a demo, the project has no quality control. Software has tests; AI applications need evaluation sets, and building one is the first engineering task rather than the last.

  • Golden dataset first — representative inputs with correct outputs agreed by the people who own the process, built before the application
  • Automated scoring on every change — prompt edits, model upgrades and retrieval changes all re-scored, so improvement is demonstrated rather than asserted
  • Adversarial cases included — the ambiguous, the out-of-scope and the deliberately awkward, because production supplies these whether or not you planned for them
  • Arabic evaluated separately — English performance says little about Arabic performance, particularly in retrieval
  • Human review on a sample — automated scoring catches regressions; people catch the failures nobody thought to score
  • Live monitoring after launch — refusal rates, retrieval misses, latency and cost tracked as operational metrics, not curiosities

This is also the honest answer to the hallucination question. It is not solved by a better prompt. It is managed by grounding answers in retrieved sources, validating structure, designing refusal, routing consequential actions to human approval, and measuring continuously.

06 — Cost and latency

The two budgets nobody sets early enough

AI applications have a marginal cost per use, which is unfamiliar to organisations accustomed to licensing software once. A feature that is delightful at a hundred users a day can be indefensible at a hundred thousand.

Both budgets belong in the design conversation. What is the acceptable cost per interaction, and the acceptable wait? Those two numbers drive model selection, routing, caching and how much retrieval you can afford per query — and they are far cheaper to answer at design time than to retrofit after launch.

  • Route by task — classification and extraction to small models, reasoning to larger ones, with the split measured rather than assumed
  • Cache aggressively — repeated questions are common in enterprise use, and a cache hit costs nothing
  • Constrain context — sending an entire document because it is easier is a recurring and avoidable cost
  • Set a latency budget per surface — a chat interface tolerates two seconds; a checkout step does not
  • Instrument cost per feature — so the expensive feature is a decision rather than a surprise in the invoice
07 — Delivery

Twelve weeks to something real

  1. Scoping and data audit

    Which pattern this is, whether the data supports it, and what "correct" means for this task. Ends with a build/buy recommendation, including the recommendation not to build.

    2 weeks
  2. Evaluation set

    Representative inputs and agreed correct outputs, built with the people who own the process. Everything afterwards is measured against this.

    1 week
  3. Core build

    Retrieval, model layer, validation, guardrails and integration. Scored against the evaluation set continuously rather than at the end.

    5–6 weeks
  4. Hardening

    Adversarial testing, latency and cost tuning, PDPL review of data flows including prompts sent to third-party providers, and human review workflows.

    2 weeks
  5. Production release

    Deployment with tracing and monitoring in place, operator documentation, and handover to the team who will run it.

    1–2 weeks

Full enterprise applications run 3 to 6 months. The constraint is almost always data readiness and integration count — rarely model capability.

Python · FastAPI PyTorch · scikit-learn LangChain · LlamaIndex OpenAI · Anthropic · open models pgvector · Pinecone · OpenSearch Kubernetes · Docker AWS · Azure MLflow · Weights & Biases OpenTelemetry · Grafana PostgreSQL · Redis GitHub Actions
08 — Questions

Answered plainly

Should we build, buy or fine-tune?

Buy when the capability is generic and a vendor already does it well. Build when the value comes from your data or your process. Fine-tune only when prompting plus retrieval has been tried and measurably falls short — fine-tuning creates a model you now own, must evaluate, and must redo when the base model moves on.

What is RAG and when is it right?

Retrieval-augmented generation fetches relevant documents at query time and gives them to the model as context. It suits questions answered by a body of knowledge that changes: policies, contracts, product documentation, case history. Usually cheaper, more current and more auditable than fine-tuning, because you can show which source produced an answer.

How do you know an AI application is working?

Through an evaluation set built before the application: representative inputs with agreed correct outputs, scored automatically on every change. Without it, quality is whatever the last person to try it thought, and every model change is an unmeasured risk.

How do you handle hallucination?

By constraining what the model is asked to do. Grounding answers in retrieved sources with citations, refusing rather than guessing when retrieval returns nothing relevant, validating structured output against a schema, and routing consequential actions to human approval. Managed by design, not eliminated by a better prompt.

Do AI applications work in Arabic?

Yes, with deliberate handling. Modern Standard Arabic is well supported and Gulf dialect is workable. Retrieval is the usual weak point, since Arabic morphology and normalisation affect chunking and embedding more than in English. Arabic evaluation sets are built alongside English ones rather than assumed equivalent.

How long does it take?

A first production release typically takes 12 weeks including the evaluation harness. Full enterprise applications run 3 to 6 months, constrained by data readiness and integration count rather than model capability.

Are AI applications PDPL compliant?

They can be, if designed that way. Lawful basis, data residency, retention and audit logging affect architecture — and prompts sent to third-party model providers are themselves a data transfer that has to be assessed. Where policy requires it, models run inside your environment.

Start here

Bring the use case. We'll tell you if it needs AI.

A two-week scoping engagement establishes which pattern your problem fits, whether your data supports it, what "correct" means, and what it will cost per interaction at your volume — including the recommendation to buy or to do nothing.