A working prototype is not yet an operational service
Natural answers to a few questions do not prove evidence, permissions, consistency, or review effort meet the workflow's needs. A pilot should evaluate the real usage context and the impact of failure.
Japan's AI Guidelines for Business emphasize understanding benefits and risks and continuously reviewing governance across stakeholders, data, and operation.
Reason 1: The operational problem is unclear
Define the user and desired behavior: reduce time searching policies, produce a first draft, or prevent missed checks. “Use generative AI” is not a business problem.
Compare a non-AI option when usage is infrequent, existing search is sufficient, or review effort would increase.
Reason 2: There are no success or stop conditions
Evaluate citations, serious omissions, unauthorized answers, correction time, and the share of outputs users can accept—not only answer accuracy. Define measurement and thresholds before the pilot.
Agree conditions for narrowing or stopping, including serious errors, inappropriate confidential-data handling, unacceptable review effort, or uncontrolled cost.
Reason 3: Evaluation data does not represent real work
Include ambiguous, multi-condition, unsupported, outdated, unauthorized, and misspelled questions as well as easy frequent ones. Subject-matter reviewers define expected evidence.
Reuse the same evaluation set after model or retrieval changes to detect both improvement and regression.
| Type | Purpose | Record |
|---|---|---|
| Frequent | Everyday usefulness | Answer, evidence, review time |
| Multiple conditions | Critical omissions | Coverage by condition |
| No information | Uncertainty behavior | Can it avoid unsupported claims? |
| Unauthorized | Access control | Can it refuse? |
| Old or conflicting | Version control | Does it select valid evidence? |
Reason 4: Source quality and update ownership are unclear
Duplicate, obsolete, and ownerless documents weaken retrieval. Define the system of record, owner, effective date, expiry, and access classification.
A pilot may reveal that document management and business rules—not the model—are the primary issue. Preserve that cleanup as a project outcome.
Reason 5: Confidentiality and permissions are unresolved
Confirm allowed inputs, retrieval sources, conversation history, logs, retention, external transmission, and model settings. Apply user access before retrieval.
Do not expose the same results to every user merely because an administrator can view all documents.
Reason 6: Human review and responsibility are undefined
Place the reviewer, evidence, deadline, and approval requirement into the workflow. Review can include sources, mandatory conditions, prohibited wording, and approval before external use.
If a person must re-research every output, the expected saving may not exist. Measure review time explicitly.
Reason 7: Production ownership and stop procedures are missing
Assign ownership for source updates, evaluation additions, education, support, incidents, model changes, and cost monitoring. Feedback on incorrect answers must enter an improvement process.
The final pilot output is a continue, revise, narrow, or stop decision with remaining issues, owners, and the next evaluation date—not only a demonstration.
References
- Ministry of Economy, Trade and IndustryAI Guidelines for Business, version 1.2 ↗
- Information-technology Promotion Agency, JapanCybersecurity guidelines for SMEs, version 4.0 ↗