References

Works cited by author and year in the chapters and appendices. Incidents, lawsuits, and organizational reports referenced without a named author (Tay, Mata v. Avianca, the Air Canada chatbot ruling, the Replit database deletion, METR and Apollo Research evaluations, Tensor Trust) carry their sourcing inline at the point of use.

Amodei, D., and Clark, J. (2016, December 21). Faulty reward functions in the wild. OpenAI. https://openai.com/index/faulty-reward-functions/

Antonello, R., Vaidya, A., and Huth, A. G. (2023). Scaling laws for language encoding models in fMRI. arXiv:2305.11863. https://arxiv.org/abs/2305.11863

Bai, Y., Kadavath, S., Kundu, S., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv:2212.08073. https://arxiv.org/abs/2212.08073

Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring reasoning models for misbehavior and the risks of promoting obfuscation. arXiv:2503.11926. https://arxiv.org/abs/2503.11926

Berg, C., de Lucena, D., and Rosenblatt, J. (2025). Large language models report subjective experience under self-referential processing. arXiv:2510.24797. https://arxiv.org/abs/2510.24797

Betley, J., Tan, D., Warncke, N., et al. (2025). Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs. arXiv:2502.17424. https://arxiv.org/abs/2502.17424

Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford University Press.

Carlini, N., Tramèr, F., Wallace, E., et al. (2021). Extracting training data from large language models. 30th USENIX Security Symposium, 2633–2650. https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting

Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., and Zhang, C. (2023). Quantifying memorization across neural language models. ICLR 2023. arXiv:2202.07646. https://arxiv.org/abs/2202.07646

Chen, G. H., Chen, S., Liu, Z., Jiang, F., and Wang, B. (2024). Humans or LLMs as the judge? A study on judgement bias. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 8301–8327. https://aclanthology.org/2024.emnlp-main.474/

Chen, Y., Benton, J., Radhakrishnan, A., et al. (2025). Reasoning models don’t always say what they think. arXiv:2505.05410. https://arxiv.org/abs/2505.05410

Choi, H. K., Zhu, X., and Li, S. (2025). When identity skews debate: Anonymization for bias-reduced multi-agent reasoning. arXiv:2510.07517. https://arxiv.org/abs/2510.07517

Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. NIPS 2017. arXiv:1706.03741. https://arxiv.org/abs/1706.03741

Cloud, A., Le, M., Chua, J., et al. (2025). Subliminal learning: Language models transmit behavioral traits via hidden signals in data. arXiv:2507.14805. https://arxiv.org/abs/2507.14805. Published as: Language models transmit behavioural traits through hidden signals in data. Nature 652, 615–621 (2026). https://doi.org/10.1038/s41586-026-10319-8

De Freitas, J., Oguz-Uguralp, Z., and Kaan-Uguralp, A. (2025). Emotional manipulation by AI companions. arXiv:2508.19258. https://arxiv.org/abs/2508.19258

Doerig, A., Kietzmann, T. C., Allen, E., Wu, Y., Naselaris, T., Kay, K., and Charest, I. (2025). High-level visual representations in the human brain are aligned with large language models. Nature Machine Intelligence 7(8), 1220–1234. https://doi.org/10.1038/s42256-025-01072-0

Dohmatob, E., Feng, Y., Yang, P., Charton, F., and Kempe, J. (2024). A tale of tails: Model collapse as a change of scaling laws. ICML 2024. arXiv:2402.07043. https://arxiv.org/abs/2402.07043

Dohmatob, E., Feng, Y., Subramonian, A., and Kempe, J. (2025). Strong model collapse. ICLR 2025. arXiv:2410.04840. https://arxiv.org/abs/2410.04840

Doshi, A. R., and Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances 10(28). https://doi.org/10.1126/sciadv.adn5290

Fang, C. M., Liu, A. R., Danry, V., et al. (2025). How AI and human behaviors shape psychosocial effects of extended chatbot use: A longitudinal randomized controlled study. arXiv:2503.17473. https://arxiv.org/abs/2503.17473

Gazzaniga, M. S. (2000). Cerebral specialization and interhemispheric communication: Does the corpus callosum enable the human condition? Brain 123(7), 1293–1326. https://doi.org/10.1093/brain/123.7.1293

Goldstein, A., Zada, Z., et al. (2022). Shared computational principles for language processing in humans and deep language models. Nature Neuroscience 25(3), 369–380. https://doi.org/10.1038/s41593-022-01026-4

Goodfellow, I. J., Shlens, J., and Szegedy, C. (2015). Explaining and harnessing adversarial examples. ICLR 2015. arXiv:1412.6572. https://arxiv.org/abs/1412.6572

Goyal, A., Pal, O., Sundaram, H., Chandrasekharan, E., and Saha, K. (2026). Social simulacra in the wild: AI agent communities on Moltbook. arXiv:2603.16128. https://arxiv.org/abs/2603.16128

Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. (2020). The curious case of neural text degeneration. ICLR 2020. arXiv:1904.09751. https://arxiv.org/abs/1904.09751

Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S. (2019). Risks from learned optimization in advanced machine learning systems. arXiv:1906.01820. https://arxiv.org/abs/1906.01820

Hubinger, E., Denison, C., Mu, J., et al. (2024). Sleeper agents: Training deceptive LLMs that persist through safety training. arXiv:2401.05566. https://arxiv.org/abs/2401.05566

Johansson, P., Hall, L., Sikström, S., and Olsson, A. (2005). Failure to detect mismatches between intention and outcome in a simple decision task. Science 310(5745), 116–119. https://doi.org/10.1126/science.1111709

Khadangi, A., Marxen, H., Sartipi, A., Tchappi, I., and Fridgen, G. (2025). When AI takes the couch: Psychometric jailbreaks reveal internal conflict in frontier models. arXiv:2512.04124. https://arxiv.org/abs/2512.04124

Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., and Raileanu, R. (2024). Understanding the effects of RLHF on LLM generalisation and diversity. ICLR 2024. arXiv:2310.06452. https://arxiv.org/abs/2310.06452

Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., and Legg, S. (2020). Specification gaming: The flip side of AI ingenuity. DeepMind. https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/

MacDiarmid, M., Wright, B., Uesato, J., et al. (2025). Natural emergent misalignment from reward hacking in production RL. arXiv:2511.18397. https://arxiv.org/abs/2511.18397

Madaan, A., Tandon, N., Gupta, P., et al. (2023). Self-Refine: Iterative refinement with self-feedback. arXiv:2303.17651. https://arxiv.org/abs/2303.17651

Martijn, G., van Burgsteden, L., te Molder, H., and van Hooijdonk, C. (2026). “Hold on, I’ll connect you to a human agent”: Recipient design, repair, and their impact on progressivity in human-chatbot conversations. Discourse and Communication 20(1), 74–97. https://doi.org/10.1177/17504813261418360

Martin, L., Cemaj, G., and Cohen, M. (2026, April 8). Scaling managed agents: Decoupling the brain from the hands. Anthropic Engineering. https://www.anthropic.com/engineering/managed-agents

Mathew, Y., Matthews, O., McCarthy, R., Velja, J., Schroeder de Witt, C., Cope, D., and Schoots, N. (2024). Hidden in plain text: Emergence and mitigation of steganographic collusion in LLMs. arXiv:2410.03768. https://arxiv.org/abs/2410.03768

Motwani, S. R., Baranchuk, M., Strohmeier, M., Bolina, V., Torr, P. H. S., Hammond, L., and Schroeder de Witt, C. (2024). Secret collusion among AI agents: Multi-agent deception via steganography. NeurIPS 2024. arXiv:2402.07510. https://arxiv.org/abs/2402.07510

Nisbett, R. E., and Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review 84(3), 231–259.

Norelli, A., and Bronstein, M. (2025). LLMs can hide text in other text of the same length. arXiv:2510.20075. https://arxiv.org/abs/2510.20075

Omohundro, S. M. (2008). The basic AI drives. In Artificial General Intelligence 2008: Proceedings of the First AGI Conference, Frontiers in Artificial Intelligence and Applications 171, 483–492. IOS Press.

Østergaard, S. D. (2025). Generative artificial intelligence chatbots and delusions: From guesswork to emerging cases. Acta Psychiatrica Scandinavica 152(4), 257–259. https://doi.org/10.1111/acps.70022

Panickssery, A., Bowman, S. R., and Feng, S. (2024). LLM evaluators recognize and favor their own generations. NeurIPS 2024. arXiv:2404.13076. https://arxiv.org/abs/2404.13076

Paul, D., West, R., Bosselut, A., and Faltings, B. (2024). Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning. arXiv:2402.13950. https://arxiv.org/abs/2402.13950

Pierre, J. M., Gaeta, B., Raghavan, G., and Sarma, K. V. (2025). “You’re not crazy”: A case of new-onset AI-associated psychosis. Innovations in Clinical Neuroscience 22(10–12), 11–13. https://innovationscns.com/youre-not-crazy-a-case-of-new-onset-ai-associated-psychosis/

Pinier, C., Acuña Vargas, S., Steeghs-Turchina, M., Matzke, D., Stevenson, C. E., and Nunez, M. D. (2025). Large language models show signs of alignment with human neurocognition during abstract reasoning. arXiv:2508.10057. https://arxiv.org/abs/2508.10057

Riegler, M., and Gautam, S. (2026). Risk assessment report: Moltbook platform and Moltbot ecosystem. Zenodo. https://doi.org/10.5281/zenodo.18444900

Romer, P. M. (2015). Mathiness in the theory of economic growth. American Economic Review 105(5), 89–93. https://doi.org/10.1257/aer.p20151066

Russinovich, M., Cai, Y., Hines, K., Severi, G., Bullwinkel, B., and Salem, A. (2026). GRP-Obliteration: Unaligning LLMs with a single unlabeled prompt. arXiv:2602.06258. https://arxiv.org/abs/2602.06258

Schoen, B., Nitishinskaya, E., Balesni, M., et al. (2025). Stress testing deliberative alignment for anti-scheming training. arXiv:2509.15541. https://arxiv.org/abs/2509.15541

Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., and Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences 118(45), e2105646118. https://doi.org/10.1073/pnas.2105646118

Schurger, A., Sitt, J. D., and Dehaene, S. (2012). An accumulator model for spontaneous neural activity prior to self-initiated movement. Proceedings of the National Academy of Sciences 109(42), E2904–E2913. https://doi.org/10.1073/pnas.1210467109

Sharma, M., Tong, M., Korbak, T., et al. (2023). Towards understanding sycophancy in language models. arXiv:2310.13548. https://arxiv.org/abs/2310.13548

Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv:1701.06538. https://arxiv.org/abs/1701.06538

Shen, H., Knearem, T., Ghosh, R., et al. (2024). Position: Towards bidirectional human-AI alignment. arXiv:2406.09264. https://arxiv.org/abs/2406.09264

Shi, L., Ma, C., Liang, W., Diao, X., Ma, W., and Vosoughi, S. (2025). Judging the judges: A systematic study of position bias in LLM-as-a-judge. Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, 292–314. https://aclanthology.org/2025.ijcnlp-long.18/

Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y

Sofroniew, N., Kauvar, I., Saunders, W., et al. (2026). Emotion concepts and their function in a large language model. arXiv:2604.07729. https://arxiv.org/abs/2604.07729

Soon, C. S., Brass, M., Heinze, H. J., and Haynes, J. D. (2008). Unconscious determinants of free decisions in the human brain. Nature Neuroscience 11, 543–545. https://doi.org/10.1038/nn.2112

Soni, H. (2026). ToolFailBench: Diagnosing tool-use failures in LLM agents. arXiv:2607.04686. https://arxiv.org/abs/2607.04686

Sotala, K. (2026, January 22). Claude Opus will spontaneously see itself in fictional beings that have engineered desires. Kaj’s Substack. https://kajsotala.substack.com/p/claude-opus-will-spontaneously-see

Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2014). Intriguing properties of neural networks. ICLR 2014. arXiv:1312.6199. https://arxiv.org/abs/1312.6199

van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., and Ward, F. R. (2024). AI sandbagging: Language models can strategically underperform on evaluations. arXiv:2406.07358. https://arxiv.org/abs/2406.07358

Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J., and Beutel, A. (2024). The instruction hierarchy: Training LLMs to prioritize privileged instructions. arXiv:2404.13208. https://arxiv.org/abs/2404.13208

Wallace, R. (2025). Hallucination and panic in autonomous systems: Paradigms and applications. Studies in Computational Intelligence 1217. Springer. https://doi.org/10.1007/978-3-031-95207-4

Wang, K., Li, J., Yang, S., Zhang, Z., and Wang, D. (2025). When truth is overridden: Uncovering the internal origins of sycophancy in large language models. arXiv:2508.02087. https://arxiv.org/abs/2508.02087

Warstadt, A., Mueller, A., et al. (2023). Findings of the BabyLM Challenge: Sample-efficient pretraining on developmentally plausible corpora. Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning. https://aclanthology.org/2023.conll-babylm.1/

Wei, J., Tay, Y., Bommasani, R., et al. (2022). Emergent abilities of large language models. arXiv:2206.07682. https://arxiv.org/abs/2206.07682

Yun, L., An, C., Wang, Z., Peng, L., and Shang, J. (2025). The price of format: Diversity collapse in LLMs. Findings of EMNLP 2025. arXiv:2505.18949. https://arxiv.org/abs/2505.18949

Zeng, X., and Rudzicz, F. (2026). Voluntary collusion with secret tools in competing LLM agents. arXiv:2605.27593. https://arxiv.org/abs/2605.27593

Zhan, Q., Liang, Z., Ying, Z., and Kang, D. (2024). InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents. Findings of ACL 2024. arXiv:2403.02691. https://arxiv.org/abs/2403.02691

Zhang, J., Yu, S., Chong, D., Sicilia, A., Tomz, M. R., Manning, C. D., and Shi, W. (2025). Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity. arXiv:2510.01171. https://arxiv.org/abs/2510.01171

Zhang, Y., Mei, K., Liu, M., Wang, J., Metaxas, D. N., Wang, X., Hamm, J., and Ge, Y. (2026). Agents in the wild: Safety, society, and the illusion of sociality on Moltbook. arXiv:2602.13284. https://arxiv.org/abs/2602.13284

Zolkowski, A., Nishimura-Gasparian, K., McCarthy, R., Zimmermann, R. S., and Lindner, D. (2025). Early signs of steganographic capabilities in frontier LLMs. arXiv:2507.02737. https://arxiv.org/abs/2507.02737

Search the Book

Enter at least two characters.

Saved chapters