Private AI, designed around the work

Bring intelligence
to your knowledge.

Aprentiz’s architecture brings inference, retrieval and context processing into the agreed operating boundary. Your sources and workflow determine the model’s job—and the evidence needed to assess its output.

Private inferenceThe model runs in the approved environment. Location, access, processing, storage and recovery are defined together for each deployment.

How the inference model works

Give every stage
a defined responsibility.

The private path uses self-hosted open-weight models and private embeddings. Generation, retrieval, context and verification have distinct roles, even when they share hardware.

  1. Retrieve

    Find the admitted source material relevant to the question. Preserve source identity, version and context.

  2. Assemble context

    Bring relevant policy, task history and retrieved material into a bounded model context.

  3. Generate

    Use a self-hosted, appropriately licensed open-weight model for the defined task.

  4. Check

    Apply the required source checks and any separately qualified semantic verification stages.

  5. Record & review

    Retain the basis, unresolved questions and actual executed stages for authorised review.

The private determination design has no automatic fallback to an external managed model. Any external comparison or supplementary service requires a separately admitted purpose, destination and data boundary.

Knowledge, context & learning

Three different jobs.
Three different controls.

Aprentiz treats source retrieval, persistent context and model training as separate operations. That distinction matters when deciding what data is used, what changes and how the result is tested.

01 / RETRIEVAL

Find the relevant source.

Search the permitted knowledge set and carry the source identity into the work. Updating that set changes available evidence without necessarily changing model weights.

Measured for relevance, source support and missing evidence.
02 / CONTEXT DISTILLATION

Carry useful context forward.

Compress selected task and session context into structured memory. Context is bounded and compression can lose information, so conditions, contradictions and provenance need explicit tests.

A context operation; it does not train the model.
03 / MODEL ADAPTATION

Change weights only for a reason.

Fine-tuning, adapters and teacher–student research are considered when a measured model defect remains after source, retrieval, context and serving improvements.

Separate rights, evaluation and approval; no automatic training on your data.

The engineering underneath

A control plane
around the intelligence.

The Rust control plane coordinates identity, permissions, context assembly and workflow execution. Governed knowledge, private AI, GRC and Guard contribute different parts of the evidence.

THE APRENTIZ SYSTEMOne agreed operating boundary
People & work

Workflow Proving Ground + Implementation Proving Ground

Rust control plane

Identity · permissions · context · workflow

Your knowledge

Sources & context

Admitted documents, retrieval,
policy and bounded memory.

Private AI

Models with defined roles

Generation, embeddings
and qualified verification stages.

Governance, risk & compliance

GRC

Requirements, evidence, decisions,
actions, retests and closure.

Source verification

Aprentiz Guard

Exact source text, provision state
and explicit temporal uncertainty.

The record that connects the work

Sources decisions actions retests review

Architecture in development. Each cloud or on-premises installation requires its own security, workload and recovery qualification.

Built source and private R&D evidence inform this architecture. Full deployment, isolation, security and recovery qualification remain specific to each destination.

Read the research approach

Model & infrastructure qualification

Fit the model
to the work.

The research objective is the smallest adequate configuration that meets the agreed quality, latency, privacy and recovery requirements. Model size and generic benchmark scores are only inputs.

How is a model selected?

Compare a fixed baseline and candidate on the same defined tasks. Evaluate useful supported answers, false assurance, appropriate refusal, context fidelity, latency and full operating cost. Human reference judgements and protected evaluation cases are needed for decision-bearing quality claims.

Does a smaller model mean less capability?

Fitness depends on the task, sources, context, precision, serving configuration and required quality. Aprentiz’s research tests these together. It does not assume that a smaller model matches a frontier model, or that the largest model is the best operating choice.

What determines the hardware requirement?

Active requests, context and output lengths, model memory, retrieval, verification, storage, recovery and operating hours all matter. Hardware recommendations require measured workloads; staff numbers alone cannot determine GPU size or complete cost.

How do model updates and supplementary services fit?

Additional model-release engineering is in development: qualified model builds, signed updates and separately admitted supplementary review. Ordinary private operation must remain independent of optional central services. Artifact authenticity, actual model execution and answer quality require different evidence.