Field notes & Wiki.
Reference architectures, engineering benchmarks, and the thinking behind the work.
The Air-Gapped GPU Fallacy: Realities of Running Offline CUDA Clusters
Unplugging the ethernet cable doesn't magically secure your GPU cluster. How dependency mirrors, PyPI wheels, CUDA driver updates, and model weights introduce supply-chain vulnerabilities in offline enterprise environments.
Stop Filtering Prompts: Why Autonomous AI Agents Require System-Call Sandboxing
Relying on prompt filters to stop agent hijacking is a security placebo. If an agent holds broad API credentials, an attacker needs only one indirect injection to exfiltrate your database. Here is how to sandbox tool execution at the runtime layer.
Dynamic Model Routing: Cutting LLM Spend Without Cutting Capability
Most enterprise LLM spend is not a model-selection problem, it is a routing problem. A reference architecture for a governed AI gateway, and the four places it breaks.
Vector Access Control Leaks: Why Post-Retrieval Filtering is a Security Lie
An agent that can be talked into querying your database is only as contained as the credentials you handed it. How embedding inversion attacks leak raw data, and why RBAC must be enforced at the storage engine level.
Post-Quantum Cryptographic Readiness for Legacy Infrastructure
Post-quantum migration is an inventory problem before it is a cryptography problem. Using Mosca's inequality to decide what is actually urgent, and why crypto-agility is the real deliverable.
New Field Notes in your inbox.
We publish when we have something worth saying — reference architectures, benchmark tests, and engineering analysis. No cadence, no spam.