Open weights are an opportunity, not a boundary
The defensible boundary is the verified, continuously tested system around the model — not the fact that inference runs on your hardware.
“Local model” is not “contained system”
Leaders hear local and infer private. In practice, the model sits inside a larger loading, identity, network, data, and tool path.
What the buyer sees
What actually runs
The hidden data path
Nine linked paths all sit inside the unit leadership is actually approving.
Nine live exposure categories
The risk is not theoretical; it follows the data and authority already present in the workload.
Prompts, caches, indexes, logs, and remote routes.
Environment, home directories, metadata, and agent forwarding.
Retrieval, telemetry, backups, and residency drift.
Classification and cross-tenant retrieval failures.
Verbatim server and observability logs.
Commands, tickets, database rows, and browser content.
Semantic leakage and membership inference.
Swap, backups, dumps, caches, and graphics processing unit (GPU) memory.
Unvalidated recommendations gaining authority.
Retention extends beyond the intended session
Content inherits the retention posture of every store it touches.
An unbounded model becomes an unbounded operating decision
The loss shows up as legal exposure, silent retention, tool-authority blast radius, and decisions no one validated.
Procurement exposure
Custom terms, revenue triggers, gated distribution, and unverifiable lineage.
Silent retention
Prompts and records entering systems not scoped for sensitive content.
Authority misuse
Model recommendations executed with a broker's or operator's privilege.
Poisoned decisions
Hostile content or stale permissions shaping a trusted workflow.
Eight layers, each with an owner and evidence
The model process lives inside the boundary. It does not enforce the boundary.
Seven isolation tiers
Data sensitivity, provenance uncertainty, tool authority, tenancy, regulation, criticality, and blast radius set the minimum.
Supply-chain promotion gate
A model becomes approved only after its complete loading stack produces verifiable evidence.
Tools belong behind an authenticated broker
Validate identity, schema, arguments, path, tenant, amount, retry, time, and human approval before execution.
Broker validates
Schema · canonical path · tenant · web address · query scope · amount · retries · expiry
Broker never does
Execute concatenated strings, inherit model authority, or reuse an approval after arguments change.
Data, retrieval, and retention follow classification
Classification travels with the document; permissions apply on every retrieval; deletion propagates through embeddings, memory, logs, and backups.
| Class | Prompts and responses | Events | Embeddings and memory |
|---|---|---|---|
| Public | Optional content log | Standard event schema | Retained, user-scoped |
| Internal | No content by default | IDs, model, tokens, latency, tool decision | Source and tenant tagged |
| Restricted | Disabled; isolated forensic exception | Redacted metadata and evidence reference | Tenant-filtered; memory by approval |
| Regulated | Disabled unless obligation requires it | Minimum required, residency tagged | Residency-bound; memory disabled |
macOS, Windows, and Linux enforce the same intent differently
Dedicated identity, process confinement, read-only artifacts, denied egress, minimized telemetry, and tested cleanup remain constant.
macOS
- launchd and dedicated users
- App Sandbox and Hardened Runtime
- Network Extension and Endpoint Security
- Virtualization.framework for stronger isolation
Windows
- Service security identifier and restricted identity
- AppContainer and Job Objects
- App Control and Defender Firewall
- Virtualization-based security, Windows Sandbox, and Hyper-V
Linux
- systemd sandboxing and cgroups v2
- Namespaces, seccomp, capabilities
- AppArmor or SELinux and Landlock
- nftables and rootless containers
Monitoring without alert fatigue
Suppression can group duplicate notifications. It must never suppress event collection or evidence.
Keep native meters separate before any planning bridge
GPU-hours, power, storage, labor, monitoring, review, downtime, and audit effort are different costs. Avoided breach is not guaranteed savings.
Low capital, endpoint burden
Utilization drives economics
Deallocation discipline matters
Tenancy and on-call dominate
Duplication, audit, and staffing
Thirty days to scale, revise, contain, or stop
The pilot produces inventory, containment, evidence, one tested workload, and a written residual-risk decision.
Inventory, immutable identifiers, contain the highest-risk artifact.
Intake, provenance, isolated load, signed registry promotion.
Chosen isolation tier, gateway, telemetry, canaries.
Full test, kill switch, revocation, rollback, rebuild, decision.
Six decisions belong to leadership
These decisions set the perimeter within which implementation teams can move quickly without improvising risk.
The pilot exits through a scorecard, not a feeling
Each measure has a target, actual, evidence reference, owner, and direct effect on the scale decision.
| Measure | Target | Status | Decision effect |
|---|---|---|---|
| Artifacts with immutable identifiers and approval | 100% | Stop promotion if below | |
| Network and metadata tests passed | 100% | Contain immediately on failure | |
| Mediated, attributable tool calls | 100% | Disable tool on failure | |
| Successful canary exfiltration or persistence | 0 | Stop and investigate | |
| Kill switch, revocation, rollback, rebuild | All complete | No scale without completion |
Approve the perimeter, not the intuition
The full report supplies the architecture, platform field manuals, fourteen-scenario control matrix, sixty failure modes, copyable policies, test protocol, incident playbooks, economics, and phased rollout.
Harden the path
Open weights create control opportunities. Provenance, isolation, mediation, evidence, and recovery turn them into assurance.
Choose the first workload
Inventory it, contain it, test it, and record a scale-or-stop decision within thirty days.