# Open questions and evidence gaps

**15 September 2026**

## Comparable outcomes

How much does deception improve detection in a given scenario? Studies differ in metrics, placement and adversaries. The [NCSC](https://www.ncsc.gov.uk/blog-post/cyber-deception-trials-what-weve-learned-so-far) found missing outcome measures. A useful comparison should publish its environment, simulated threat, baseline, alert timing, response actions and maintenance cost.

## Realism and decoy detection

How does honeypot value change when an adversary recognizes artificial signals? The catalog includes work such as [Towards Systematic Honeytoken Fingerprinting](https://doi.org/10.1145/3433174.3433599), initially as metadata only. Its methods must be read and linked to documented countermeasures before drawing conclusions.

## Cloud, identity and industrial systems

Which artifacts are safe and observable in cloud, identity, OT and IoT? Field reports need configuration, system limits, operational consequences and negative outcomes. The [NCSC trials](https://www.ncsc.gov.uk/blog-post/cyber-deception-trials-what-weve-learned-so-far) covered multiple environments and identified configuration risks.

## Language models

What interaction depth justifies the cost of an AI-based honeypot? How can inconsistent responses or abuse of the decoy as attack infrastructure be controlled? [LLM Honeypot](https://arxiv.org/abs/2409.08234) and [HoneyGPT](https://arxiv.org/abs/2406.01882) report specific evaluations. Comparable-condition replication is needed.

## Field experience and worldwide coverage

We need stories describing failed deployments, maintenance and ignored alerts. Verifiable products, services and experiences outside well-documented English-language markets are also underrepresented. The Atlas will publish its searches and gaps by language and region to guide contributions.
