© 2026 Medoya LLC
Powered by voilamia
Subscribe to receive new insights
Medoya Labs
A proof-of-concept reference architecture I presented at ISC2 Security Congress 2025: an API and AI gateway in front, a policy engine that scores the risk of every request, and a second model that checks the first. Built on open source and run on premises, for imaging workflows where patient data cannot leave the building.
In early 2025, a study in Nature Communications tested vision-language models on medical images carrying hidden instructions: text a radiologist would not notice, such as dark text on a dark background inside a scan. Across 594 attacks on four models, every model could be led to miss lesions it would otherwise report. GPT-4o followed the hidden instruction in 67% of its 162 attacks.
That is a semantic attack. The image is valid and the request is well formed, so a web application firewall or a data loss prevention filter has nothing to match. The defense has to understand what the content means, and it has to run on every request without slowing a clinical workflow to a crawl.
MedX Gateway is the reference architecture I built to answer that, and the Twin Gatekeeper is its most distinctive piece. Every request to a model, whether it comes from a chatbot, a clinician's tool or an AI agent, passes through three layers: access, governance and guardrails. Each layer has one job, and each is an open-source component an organization can run on its own hardware.
Layer 1, access: Kong AI Gateway. Open-source Kong sits in front of every model. It authenticates the caller, applies rate limits and quotas, validates the request, routes it to the right model provider, holds the provider credentials in one place, counts tokens, and opens the trace every later step writes into. I also evaluated Portkey, Gloo Gateway, Tyk, LiteLLM and TrueFoundry for this layer.
Layer 2, governance: Open Policy Agent. Policies live as code in Rego, so they are reviewed, versioned and audited like any other code. For each request the policy engine computes a risk score from four factors: who is asking (identity), from where (environment), whether this matches their usual activity (behavior), and what they are asking for (resource: data classification, tools, attachments). The weights come from the organization's risk appetite. The score decides which Layer 3 checks run, and a request whose remaining risk is still above the threshold after those checks is blocked. The same logic runs again on the answer before it is released. The process follows the NIST AI Risk Management Framework: map, measure, manage, govern.
Layer 3, guardrails: NVIDIA NeMo Guardrails and the Twin Gatekeeper. PII and PHI detection and redaction, topic control, output validation and hallucination checks. The Twin Gatekeeper is a second model, the checker, that inspects what goes into and comes out of the main model, the generator. The checker is smaller, runs locally and has a narrower job: decide whether an input carries instructions it should not, and whether an output is safe to release. A vision checker does the same for images.
Not every request deserves every check. An IT support specialist asking about an internal system gets authentication, rate limiting and PII detection, and the answer comes back quickly. An imaging data coordinator attaching a scan from an outside imaging center is a different case: medical images are a known attack vector and external images may already be compromised, so the policy adds the vision checker and the full policy review before the generator sees anything. The live demo walked through both, ending with a prompt injection hidden in a medical image being caught before it reached the model.
Mapped against the 2025 OWASP Top 10 for LLM Applications, the architecture addresses seven of the ten risks: prompt injection, sensitive information disclosure, improper output handling, excessive agency, system prompt leakage, misinformation and unbounded consumption. Data and model poisoning is only partly covered. Supply chain vulnerabilities and vector and embedding weaknesses are not covered: they live in the build pipeline and the retrieval store rather than in the request path, and they need their own controls, such as signed models, an AI bill of materials and access control on the vector database.
The talk ended with a maturity roadmap, and it is the one I still use with clients. In the first six months, put a gateway in front of every model for visibility, logging and cost tracking, and run PII detection and topic rails in audit-only mode while you find the shadow AI. Between six and eighteen months, turn on the checker model for high-risk use cases, let the policy engine choose guardrails by risk, and send gateway logs to your SIEM. After that, use the gateway as the layer that keeps you free to change model vendors, and move to continuous red teaming.
In a readiness review we start with the questions from the talk: who can use which model, whether every prompt, response and security decision lands in an audit trail, how AI cost traces back to a person, an agent or a department, and whether regulated data is scanned before it ever leaves the network.
A proof of concept and reference architecture, presented with a live demo at ISC2 Security Congress 2025 in the session "The Twin Gatekeeper Pattern: How to Stop Semantic Threats in Multi-Modal LLMs." The demo ran on our lab hardware: the gateway, the policy engine and the checker models on premises, on an AMD EPYC server with four RTX 3090 GPUs under Proxmox, with GPT-5 on Azure as the generator. The code is not published. The products named here are the ones I chose for the build; none of their vendors sponsored or reviewed this work.
Book a free 30-minute call. Tell me which AI tools your team uses today and one workflow you'd like an agent to take on. You'll leave knowing the three biggest things in your way and a first step for each.