Source · DiggingBeagle record
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
Each support, contradiction or context label applies to a cited Claim, not to a whole Case.
Source record
Claim-level citations (2)
- supportsTrojanStego trained a language model to leak secrets inside natural-looking text: The paper reports reliable transmission of 32-bit secrets, with 87% accuracy on held-out prompts and above 97% using majority voting across three generations in its experiments.
ACL Anthology abstract, reported accuracy
- supportsTrojanStego trained a language model to leak secrets inside natural-looking text: TrojanStego fine-tuned an LLM to encode sensitive context information into natural-looking outputs through a vocabulary-partitioning steganographic scheme.
ACL Anthology abstract
Cite this record
DiggingBeagle. “TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent.” https://diggingbeagle.com/sources/trojanstego-your-language-model-can-secretly-be-a-steganographic-privacy-leaking/
Citation guidance