AI security research / concepts

Topics

Mechanisms and patterns across the research.

31 Topics · Page 2 of 2

Search & filter

On this page only.

Published concepts

Topic

MCP control-plane exposure

An MCP endpoint exposes privileged local or remote capabilities without adequate authentication, origin checks or action constraints.

6 explicit connections

Topic

Out-of-band covert channels

Communication paths that cross an intended isolation boundary through physical emissions, environmental effects, sensors, shared infrastructure or semantic carriers outside the approved data path.

16 explicit connections

Topic

Phantom squatting: hallucinated domains as an agent-facing risk

Phantom squatting is a software-supply-chain risk in which an AI system hallucinates or invents a package/domain dependency and an attacker registers the nonexistent name so later automated consumers resolve to attacker-controlled infrastructure. The risk becomes more serious when agents can install dependencies or follow generated URLs without independent verification.

2 explicit connections

Topic

Policy-attested execution

Architecture that moves authorization out of the LLM and cryptographically binds an approved intent/policy decision to exact execution bytes.

2 explicit connections

Topic

Recursive self-improvement governance gap

Recursive self-improvement is not one evidentiary claim. A narrow system can improve a search or orchestration policy while leaving its underlying model unchanged; a stronger scenario is AI automating enough AI R&D to accelerate capability development; stronger again is uncontrolled recursive capability escalation that outpaces human evaluation and intervention. The scoped sources support active expert concern, governance debate and a futurist community tendency to compress these levels into the same label. They do not demonstrate the strongest scenario. Palisade's interviews are useful for understanding beliefs inside and around frontier labs, but Palisade itself says its interviewees are not representative and are disproportionately safety-oriented, so their extinction-risk estimates are testimony rather than population statistics or incident evidence.

1 explicit connection

Topic

Reflective optical leakage

Unintended disclosure of screen content or activity through reflections captured from eyeglasses, faces or other reflective surfaces in camera views.

1 explicit connection

Topic

Shared-service isolation failure

A service reachable from nominally isolated workloads becomes an unintended cross-session or cross-environment communication path.

3 explicit connections

Topic

Undisclosed model routing

A service silently forwards user inputs to a different model or provider than the user believes they are using.

1 explicit connection