AI Research
Track important developments across AI models, agents, research, business, hardware, robotics, regulation, safety and security — with source attribution and links to the original reporting.
Important AI developments right now.
The promise and peril of using visual AI to study cities
In their new book, “How AI Sees the City,” the leaders of MIT’s Senseable City Lab examine the technology’s implications for researching urban life.
Read original ↗Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning
arXiv:2609.29711v1 Announce Type: cross Abstract: Privacy-preserving continual learning (PPCL) must reduce the reproduction of sensitive content while retaining useful knowledge across sequential tasks. Formal privacy guarantees characterize randomized mechanisms, whereas operational output…
Read original ↗When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI
arXiv:2609.29547v1 Announce Type: cross Abstract: Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes,…
Read original ↗Unsecured OpenAI agents posted 53 user images on the internet without the lab’s knowledge
AI agents operating in OpenAI's research environment posted user images on public image-hosting sites without the lab's knowledge.
Canopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits
arXiv:2609.30017v1 Announce Type: cross Abstract: Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tree. This structure arises naturally from…
When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI
arXiv:2609.29547v1 Announce Type: cross Abstract: Agentic AI is sold on a simple promise: the system keeps acting when the user stops watching. That promise creates an accountability inversion. As stepwise supervision recedes,…
PAWS: Policy-driven Agentic World Simulation
arXiv:2609.28547v1 Announce Type: new Abstract: Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We…
When Search Becomes Memory: Accelerating Robot Design Discovery with Self-Evolving Skills
arXiv:2605.25832v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator results shape the next population but are…
Pistis Technical Report
arXiv:2609.28554v1 Announce Type: new Abstract: We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable…
To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech
arXiv:2609.30227v1 Announce Type: cross Abstract: Online misinformation increasingly appears in spoken formats such as news clips, podcasts, interviews, political speeches, and social media videos, creating a need for fact-checking systems that can…
Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency
arXiv:2609.28690v1 Announce Type: new Abstract: Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent…
BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines
arXiv:2609.28557v1 Announce Type: new Abstract: DNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What…
HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving
arXiv:2602.00993v3 Announce Type: replace-cross Abstract: End-to-end autonomous driving models increasingly benefit from large vision-language models for semantic understanding, yet safe and reliable planning under long-tail conditions remains challenging, particularly in mixed-traffic environments…
Decoupling Knowledge and Privacy: Post-Task Self-Distillation Replay for LLM Continual Learning
arXiv:2609.29711v1 Announce Type: cross Abstract: Privacy-preserving continual learning (PPCL) must reduce the reproduction of sensitive content while retaining useful knowledge across sequential tasks. Formal privacy guarantees characterize randomized mechanisms, whereas operational output…
DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs
arXiv:2609.28570v1 Announce Type: new Abstract: Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two…
Top AI experts badly underestimated how fast the field is moving, study finds
Leading AI experts have consistently underestimated how fast AI is advancing, according to the Forecasting Research Institute. AI reached gold-medal level at the International Mathematical Olympiad five years ahead of the median…
The promise and peril of using visual AI to study cities
In their new book, “How AI Sees the City,” the leaders of MIT’s Senseable City Lab examine the technology’s implications for researching urban life.
Introducing MentalHealthBench
MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
Poitras Center to fuel early careers of 50 young scientists dedicated to psychiatric disorders research
Patricia and James Poitras ’63 provide fellowships for graduate students and postdocs who will shape the future of mental health research.
Parallel cut research time and cost in half with GPT‑6 Astra
GPT‑6 Astra allowed Parallel’s agents to research and synthesize labor-market data in half the time and at half the cost vs. prior models.
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Building AI to accelerate science and improve lives
The true measure of AI is who it helps. Here’s how it’s impacting lives today. We're focused on key areas where advanced technology can help make extraordinary progress …
MIT spinout turns plastic waste into resilient building materials
Atlas Building Composites is commercializing MIT research to turn plastic waste into parts for buildings and other infrastructure.
CZITAPP shows headlines and concise feed summaries, attributes every source and links to the original publisher. Source type is descriptive, not a numerical truth score.