01Define the real job
Name the question, business scope, useful outcome and responsible professional. Make exceptions and limits explicit.
The Aprentiz Proving Ground
Two connected environments in development: one to establish how a professional workflow should operate, and another to test how its implementation will deliver that contract.
The purposeTurn a promising AI experiment into a defined workflow, measurable acceptance criteria and a reviewable implementation decision.
01 / Workflow Proving Ground
Practitioners bring the context, difficult exceptions and judgement. Aprentiz is designed to make the process, evidence and decision boundaries explicit and testable.
Name the question, business scope, useful outcome and responsible professional. Make exceptions and limits explicit.
Identify the source versions and organisational material permitted for the work. Record what is missing or excluded.
Exercise the analysis, evidence review, human decisions and failure paths. Compare the proposed process with the current way of working.
Retain the stages, roles, evidence requirements, acceptance cases and unresolved questions that implementation must preserve.
Defined inputs. Named roles. Evidence standards. Failure states. Acceptance scenarios. A clear account of what remains unresolved.
02 / Implementation Proving Ground
The technical implementation inherits the workflow’s requirements. The design keeps the original need connected to each integration, control, test and later change.
Map each workflow need to a system behaviour, data boundary, responsible owner and test.
Define integrations, model roles, controls and recovery. Make the smallest change needed to test the chosen question.
Test evidence quality, identity, isolation, missing dependencies, hostile inputs, interruption and recovery.
Bring the results, limitations and operational requirements to the authorised people. A successful demonstration is one piece of that decision.
Traceable requirements, integration maps, control tests, qualification results and the limitations an authorised release decision must consider.
The AI research underneath
Model evaluation supports both Proving Grounds. The research question is which configuration can perform the defined work usefully, at an acceptable complete cost, within the required controls.
Test source support, wrong or missing citations, contradictions, appropriate refusal and loss of conditions during context compression.
Measure latency, active requests, context limits, memory, recovery and the full cost of a useful completed workflow.
Freeze the baseline and evaluation criteria. Retain negative results. Improve, narrow, keep the baseline or stop when the evidence calls for it.
Research thresholds are agreed before testing. A model score, agreement between models or a small error-free test set does not establish professional correctness or readiness for release.