Operix.
Menu

OPERIX / INSIGHTS

Seven Reasons Generative AI Pilots Do Not Reach Operations

Separate demo quality from operational quality and define the evidence, information controls, review effort, ownership, and stop procedure needed for a production decision.

A working prototype is not yet an operational service

Natural answers to a few questions do not prove evidence, permissions, consistency, or review effort meet the workflow's needs. A pilot should evaluate the real usage context and the impact of failure.

Japan's AI Guidelines for Business emphasize understanding benefits and risks and continuously reviewing governance across stakeholders, data, and operation.

Reason 1: The operational problem is unclear

Define the user and desired behavior: reduce time searching policies, produce a first draft, or prevent missed checks. “Use generative AI” is not a business problem.

Compare a non-AI option when usage is infrequent, existing search is sufficient, or review effort would increase.

Reason 2: There are no success or stop conditions

Evaluate citations, serious omissions, unauthorized answers, correction time, and the share of outputs users can accept—not only answer accuracy. Define measurement and thresholds before the pilot.

Agree conditions for narrowing or stopping, including serious errors, inappropriate confidential-data handling, unacceptable review effort, or uncontrolled cost.

Reason 3: Evaluation data does not represent real work

Include ambiguous, multi-condition, unsupported, outdated, unauthorized, and misspelled questions as well as easy frequent ones. Subject-matter reviewers define expected evidence.

Reuse the same evaluation set after model or retrieval changes to detect both improvement and regression.

Questions to include in a pilot evaluation set
TypePurposeRecord
FrequentEveryday usefulnessAnswer, evidence, review time
Multiple conditionsCritical omissionsCoverage by condition
No informationUncertainty behaviorCan it avoid unsupported claims?
UnauthorizedAccess controlCan it refuse?
Old or conflictingVersion controlDoes it select valid evidence?

Reason 4: Source quality and update ownership are unclear

Duplicate, obsolete, and ownerless documents weaken retrieval. Define the system of record, owner, effective date, expiry, and access classification.

A pilot may reveal that document management and business rules—not the model—are the primary issue. Preserve that cleanup as a project outcome.

Reason 5: Confidentiality and permissions are unresolved

Confirm allowed inputs, retrieval sources, conversation history, logs, retention, external transmission, and model settings. Apply user access before retrieval.

Do not expose the same results to every user merely because an administrator can view all documents.

Reason 6: Human review and responsibility are undefined

Place the reviewer, evidence, deadline, and approval requirement into the workflow. Review can include sources, mandatory conditions, prohibited wording, and approval before external use.

If a person must re-research every output, the expected saving may not exist. Measure review time explicitly.

Reason 7: Production ownership and stop procedures are missing

Assign ownership for source updates, evaluation additions, education, support, incidents, model changes, and cost monitoring. Feedback on incorrect answers must enter an improvement process.

The final pilot output is a continue, revise, narrow, or stop decision with remaining issues, owners, and the next evaluation date—not only a demonstration.

References

LET’S TALK

Start with the task that is taking too much time.

Tell us about the current process, the tools you use, and what you want to change. We can work out the next step together.

Discuss your needs