Context Engineering Evals Remind Me of Infrastructure Testing
A comparison between infrastructure testing and agent evaluation, focused on hidden state, behavioral contracts, and feedback loops.
A comparison between infrastructure testing and agent evaluation, focused on hidden state, behavioral contracts, and feedback loops.
How a personal LLM wiki can separate raw evidence, curated memory, context packs, privacy boundaries, and evaluation loops.
A systems-through-line from Chef, infrastructure automation, and TDD in operations to context engineering for agentic AI.
Why reusable prompts and context packs deserve versioning, review, privacy boundaries, failure modes, and evaluation loops.