Note
The Austin ML Journal Club concluded in August 2026. This list is kept as a record and is no longer accepting suggestions. See the closing announcement.
This page collects papers, books, articles, videos, and other resources that community members suggested for future discussions. The club has concluded, so this list is now a frozen record of what was proposed but never discussed.
Suggestions at the Time of Closing
Papers the club discussed are listed on the Archives page.
- Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models (ICLR 2026)
- In-the-Flow Agentic System Optimization for Effective Planning and Tool Use (AgentFlow) (ICLR 2026)
- Position: When AI Decides Who Gets an Organ — Multi-Agentic AI Systems in Transplant Medicine Risk Amplifying Disparities
- Position: AI Capabilities Are Not Increasing Exponentially (preprint)
- Position: Beyond Reasoning Zombies — AI Reasoning Requires Process Validity (preprint)
- Position: Benchmarks Do Not Measure Deployment Readiness in Clinical AI
- SCP: Accelerating Discovery with a Global Web of Autonomous Scientific Agents
- The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
- Language models cannot reliably distinguish belief from knowledge and fact
- Enhancing Retrieval-Augmented Generation: A Study of Best Practices
- Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely
- Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
- Energy-Based Transformers are Scalable Learners and Thinkers
- Angles Don’t Lie: Unlocking Training-Efficient RL Through the Model’s Own Signals
- Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
- Text-to-LoRA: Instant Transformer Adaption
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Sample
- Petri: An open-source auditing tool to accelerate AI safety research
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
- Alignment Faking in Large Language Models
- MolmoAct: Action Reasoning Models that can Reason in Space
Previously Suggested Papers
The following were suggested by community members during our earlier meeting phases:
- Solving olympiad geometry without human demonstrations
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- TOFU: A Task of Fictitious Unlearning for LLMs
- Quantifying the impact of uninformative features on the performance of supervised classification and dimensionality reduction algorithms
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- A Mulching Proposal
- Evaluating and Mitigating Discrimination in Language Model Decisions
- Dive into Deep Learning: Coding Session #4 Attention Mechanism I (MLT Artificial Intelligence)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- MiniLLM: Large Language Models on Consumer GPUs
- The TinyLlama project
- On the Opportunities and Risks of Foundation Models
- Challenges in Deploying Machine Learning: a Survey of Case Studies
- Machine Learning and the Future of Bayesian Computation