Skip to main content
Data Products logo

Starter Services · Generative AI Pilot

One GenAI use case. Running in your business. Measured before you scale it.

Bring us a business metric worth moving. We take one defined use case from workflow to a running GenAI system, then measure quality, adoption and unit economics so you can decide whether to scale it, change it or stop.

See the use cases
From business metric to scale decision A flow from a business metric through a bounded use case, a running GenAI system, an evaluation harness and unit economics to a scale, revise or stop decision. An orange marker travels the path on a loop. THE PILOT, END TO END START HEREBusiness metric BOUND ITOne use case RUN ITGenAI system on your data, inside your controls MEASURE QUALITYEvaluation harness MEASURE ECONOMICSUnit cost at volume Scale · revise · stop, with evidence
Window2 to 4 weeks from ready-to-start kickoff
FeeFixed fee, starting at $30,000
ScopeOne use case with a defined workflow and system boundary
DecisionScale it, revise it or stop with evidence
Related delivery evidence

Selected outcomes from engagements in the same problem families. They are not performance guarantees for this pilot.

37%fewer Tier 1 live-agent contacts
28%higher first-contact resolution
15–25%of documents left for nonstandard review

Common starting points

Popular GenAI use cases, scoped down to something you can actually prove.

These are common places to start, not a menu you have to choose from. The pilot works best when the use case is tied to one business metric, one accountable owner and a bounded workflow.

Customer Service

A grounded service assistant that resolves routine questions and knows when to hand off.

Answers from your policies, product information and service history, with citations or source traceability where the workflow requires it. It can serve customers directly or support agents first, depending on risk and readiness.

The metricContainment, first-contact resolution, average handle time, or cost per resolved contact
The boundaryOne channel, one contact category, one knowledge set, one escalation path
Working meansReal or production-representative conversations, with the handoff working and the answer quality measured
The harness scoresGroundedness, resolution quality, policy adherence, escalation precision, latency and unit cost

Have another use case in mind? Good. Bring us the metric, workflow and systems involved. We will tell you whether it fits this pilot or belongs somewhere else.

What makes it more than a demo

It has to survive your data, your controls, your users and your economics.

A polished demo can prove that a model is capable of the task. The pilot is designed to prove whether the system can perform the task in your operating environment, at a quality and cost you can defend.

01

Your data and permissions

Connected to the real or production-representative source set, with the gaps, formats, entitlements and edge cases that matter after launch.

02

A real evaluation harness

A test set built from your cases and scored against the thresholds that matter for the use case, not a single generic accuracy number.

03

Inside agreed controls

Identity, access, logging, human review and escalation are designed with the people who own security, risk and the workflow.

04

Unit economics measured

Cost is measured in the unit the business actually consumes: per conversation, document, workflow run, case, asset or other agreed unit.

What "production conditions" means here

Real or production-representative data, approved integrations, representative users, agreed security controls and expected workload. A customer-facing production launch is included only when it is approved and inside the written scope.

You leave with

The system and the evidence behind the next decision.

Not a slide deck explaining what might work. A bounded system, the mechanism used to judge it, and the operating numbers required to decide what happens next.

01

A running GenAI system

Built in the agreed environment, connected to the defined data and workflow boundary, and handed over with the engagement-specific source.

02

The evaluation harness

Your test set, scoring logic and acceptance thresholds for quality, policy behavior, latency and other use-case measures.

03

Measured unit economics

The cost of operating the system at expected volume, plus the architectural and workflow levers that move that cost.

04

A scale, revise or stop decision

The evidence, tradeoffs and recommendation presented live to the people who own the next investment decision.

How it runs

Four stages inside one fixed window.

The two-to-four-week clock starts at kickoff once the access, data, environment and participant prerequisites written into the proposal are available.

STAGE 1

Bound

Fix the metric, workflow, data sources, integrations, user group, environment and acceptance thresholds in writing.

STAGE 2

Build

Connect the data, build the system and create the evaluation harness alongside it rather than at the end.

STAGE 3

Run

Put it through real or production-representative work. Measure quality, adoption, latency, exceptions and operating cost.

STAGE 4

Decide

Present what moved, what did not, what it costs, and whether the evidence supports scaling, revision or stopping.

2 to 4 weeks · fixed scope · fixed fee · from $30,000

Fit

A pilot is useful only when there is a real decision waiting at the end.

Two short lists. If you recognise yourself in the first, the pilot is the right size. If the second reads truer, we will point you at the engagement that answers your actual next question.

This is a good fit when

  • You can name the business metric or operating pain that matters.
  • There is an accountable owner who can make a scale-or-stop decision.
  • Representative data and subject-matter experts are available.
  • The workflow can be bounded to a small number of systems and users.
  • You need evidence before committing to a larger implementation.

Start somewhere else when

  • You are still asking what AI might be useful for across the organization.
  • The use case requires an enterprise-wide platform replacement before it can run.
  • No one owns the process or the metric you want to improve.
  • Representative data cannot yet be accessed or used.
  • You need a strategy, roadmap or readiness decision before a build.

Where it sits

Assessment, pilot, scale. Use the smallest engagement that answers the next question.

Three engagements, one path. Most clients enter at the step that matches what they already know, and move right only when the evidence says to.

If the use case is not settled

AI & Data Assessment Sprint

Find and score the use case, readiness gaps and dependencies before anyone starts building.

See the Assessment Sprint
If the use case is known but unproven

Generative AI Pilot

Build one bounded use case, measure it under production conditions and make the investment decision on evidence.

Start with the metric
If the value is already proven

Generative AI & AI Agents

Productionize, integrate and extend the system across additional workflows, teams or use cases.

See the practice

Before you ask

The buying questions that matter.

The seven things procurement, security and the sponsor ask before a pilot is approved, answered the way we answer them in the proposal.

What counts as production conditions?

Your real or production-representative data, approved integrations, representative users, agreed security and access controls, and an expected workload. Customer-facing production deployment is included only when it is approved and inside the agreed scope.

When does the two-to-four-week window start?

At kickoff, once the access, data, environment and participant prerequisites written into the proposal are available. That keeps an access delay from becoming a fictional delivery delay.

What does it cost?

Pilots start at $30,000. The exact fee is fixed in writing before work starts and depends on the agreed workflow, systems, integrations, user group and environment. Work inside that boundary is not billed hourly.

Which models and platforms?

We are not tied to one model or cloud. Recommendations follow your security requirements, existing technology footprint, workload, quality targets and unit economics.

What if the pilot does not hit the success threshold?

That is a valid result. You still leave with the system, harness, findings and measured economics needed to stop, revise or redirect the use case rather than scale it on optimism.

Who owns the code and evaluation assets?

The engagement is designed for handoff. You receive the engagement-specific source and evaluation assets, subject to the licensing terms of any third-party components used. Ownership details are written into the agreement before kickoff.

Will our data be used to train public models?

The answer depends on the enterprise service and contract selected. Data handling, retention and training terms are documented before build, and the architecture is configured to follow your approved controls.

Start from the business

Tell us the number you need GenAI to move.

Do not start with a model name. Give us the function, metric, workflow and systems involved. We will tell you whether it fits the pilot and, if it does, define the boundary, window and fixed fee before you commit.

FunctionWhere the work lives
MetricWhat must improve
WorkflowWhat people do today
Systems & dataWhat the system must touch

Still deciding what is worth building? Start with the AI & Data Assessment Sprint.

Have a use case? Start with the metric, workflow and systems involved.