Research
The Center’s research is anchored in two core themes: AI-First Experiences and Processes, which reimagines enterprise functions through agentic AI and next-generation user experiences that drive efficiency, intelligence, and innovation; and Secure, Responsible AI, which ensures that emerging technologies are deployed with robust security, transparency, and resilience at their core. Together, these efforts position the Center as a leading hub for shaping how enterprises design, implement, and scale AI systems that are both transformative and trustworthy.
Themes
Focus: advancing the development of AI-native systems that reimagine enterprise functions from the ground up; for example, AI agents that transform business workflows and user experiences.
Advancements in AI are transforming how humans interact with technology—moving beyond clicks and touchscreens to natural, conversational experiences through voice and free-form text. These intelligent systems can not only understand and anticipate user needs but also reimagine entire business processes by eliminating unnecessary tasks and layers. By rethinking both the user experience and the underlying operational models, organizations can unlock new forms of value, drive innovation, and achieve superior outcomes. The Center’s work in this area focuses on three key priorities: developing new interactions and experiences through personalized and intuitive interfaces; advancing process reimagining to enable streamlined workflows, intelligent decision-making, and enhanced efficiency; and leveraging data-driven insights for predictive analytics, actionable intelligence, and informed strategy development.
Focus: Security, reliability, ethical governance, and resilient AI systems; For example, ensuring that autonomous agentic solutions are robust, auditable, and aligned with enterprise values.
As generative AI systems grow more powerful and pervasive, they bring not only new opportunities but also complex security and operational challenges. Enterprises across the globe are confronting critical issues of data privacy, security, fairness, and bias—made even more urgent by the rapid evolution of AI models. Forward-looking organizations are embracing agentic modernization, and shifting away from static, legacy systems to dynamic, AI-driven architectures. The Center seeks to ensure that AI innovation advances responsibly and safely—balancing progress with accountability and foundational security principles.
Our work focuses on several key areas: strengthening defenses against novel attacks; maintenance of data integrity; building responsible AI frameworks that embed secure platform integrations into enterprise systems; and real-time identification of ongoing attacks. Together, these efforts aim to create an AI ecosystem that is secure, predictable, and trustworthy—responsible by design.
Affiliated Research Publications
The research conducted by Professors Junfeng Yang and Eugene Wu at Columbia Engineering over the past year establishes a comprehensive roadmap for the industrial-scale adoption of Agentic AI. Centering the two main themes of AI-First Experiences and Secure, Responsible AI, the following publications address the complex lifecycle of autonomous systems. First, this research explores architecting the foundational computing infrastructure, such as agentic data environments, branching benchmarks, and workflow-aware resource scheduling, necessary to support agents that can safely act across enterprise data and systems. This technical substrate is balanced by rigorous frameworks for responsible and trustworthy operation, introducing data-flow controls, compliance policies, and risk management strategies to ensure agents remain secure and compliant within regulated settings.
Agentic Data Environments — Elaine Ang, Chenxi Huang, Georgios Liargkovas, Jerry Liu, Jinhui Liu, Nikos Pagonas, Charlie Summers, Haonan Wang, Jiakai Xu, Tianle Zhou, Yusen Zhang, Zhou Yu, Zhuo Zhang, Tianyi Peng, Kostis Kaffes, Eugene Wu. IEEE Data Bulletin 2026.
Toward Systems Foundations for Agentic Exploration — Jiakai Xu, Tianle Zhou, Eugene Wu, Kostis Kaffes. SAA Workshop at SOSP 2025.
LAKEQA: An Exploratory QA Benchmark over a Million-Scale Data Lake — Haonan Wang, Jiaxiang Liu, Yurong Liu, Austin Senna Wijaya, Tianle Zhou, Eden Wu, Yijia Chen, Wanting You, Reya Vir, Daniela Pinto Veizaga, Grace Fan, Yusen Zhang, Juliana Freire, Eugene Wu. ICML 2026.
BranchBench: An Extensible Benchmark for Agentic Database Branching — Elaine Ang, Kostis Kaffes, Eugene Wu. CAIS Workshop 2026.
BranchBench: Aligning Database Branching with Agentic Demands — Elaine Ang, Sam Weldon, In Keun Kim, Kevin Durand, Kostis Kaffes, Eugene Wu. arXiv 2026.
SANA: What Matters for QA Agents over Massive Data Lakes? — Austin Senna Wijaya, Jiaxiang Liu, Haonan Wang, Eugene Wu. arXiv 2026.
Human-Data Interaction, Exploration, and Visualization in the AI Era: Challenges and Opportunities — Jean-Daniel Fekete, Yifan Hu, Dominik Moritz, Arnab Nandi, Senjuti Basu Roy, Eugene Wu, Nikos Bikakis, George Papastefanatos, Panos K. Chrysanthis, Guoliang Li, Lingyun Yu. SIGMOD Record 2026.
Data Flow Control: Data Safety Policies for AI Agents — Charlie Summers, Eugene Wu. arXiv 2026.
Please Don’t Kill My Vibe: Empowering Agents with Data Flow Control — Charlie Summers, Haneen Mohammed, Eugene Wu. CIDR 2026 Slides.
RAISE: Reliable Agent Improvement via Simulated Experience — Sahar Omidi Shayegan, Joshua Meyer, Victor Shih, Sebastian Sosa, Tianyi Peng, Kostis Kaffes, Eugene Wu, Andi Partovi, Mehdi Jamei. NeurIPS 2025 SEA Workshop.
SAGE: A Top-Down Bottom-Up Knowledge-Grounded User Simulator for Multi-turn Agent Evaluation — Ryan Shea, Yunan Lu, Liang Qiu, Zhou Yu. EACL 2026.
Detecting Privilege Escalation in Polyglot Microservices via Agentic Program Analysis — Penghui Li, Hong Yau Chong, Yinzhi Cao, Junfeng Yang. IEEE S&P 2026.
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers — Simin Chen, Jinjun Peng, Yixin He, Junfeng Yang, Baishakhi Ray. IEEE S&P 2026.
PickleBall: Secure Deserialization of Pickle-based Machine Learning Models — Andreas Kellas, Neophytos Christou, Wenxin Jiang, Penghui Li, Laurent Simon, Yaniv David, Vasileios P. Kemerlis, James C. Davis, Junfeng Yang. CCS 2025.
Do Spammers Dream of Electric Sheep? Characterizing the Prevalence of LLM-Generated Malicious Emails — Wei Hao, Van Tran, Vincent Rideout, Zixi Wang, AnMei Dasbach-Prisk, M. H. Afifi, Junfeng Yang, Ethan Katz-Bassett, Grant Ho, Asaf Cidon. IMC 2025.
Trustworthy AI Software Engineers — Aldeida Aleti, Baishakhi Ray, Rashina Hoda, Simin Chen. arXiv 2026.
The following papers, authored by other distinguished faculty and researchers at Columbia, build upon the foundational themes of AI-First Experiences and Secure & Responsible AI explored in our research tracks, offering deeper insights into the broader ecosystem of agentic systems and their application.
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving — Nikos Pagonas, Yeounoh Chung, Kostis Kaffes, Arvind Krishnamurthy. SAA Workshop at SOSP 2025.
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models — Georgios Liargkovas, Mihir Nitin Joshi, Hubertus Franke, Kostis Kaffes. arXiv 2026.
LLM Agents for Always-On Operating System Tuning — Georgios Liargkovas, Vahab Jabrayilov, Hubertus Franke, Kostis Kaffes. NeurIPS 2025.
VineLM: Trie-Based Fine-Grained Control for Agentic Workflows — Nikos Pagonas, Matthew Lou, Tianyi Peng, Dan Rubenstein, Kostis Kaffes. arXiv 2026.
Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code — Myeongsoo Kim, Dingmin Wang, Siwei Cui, Farima Farmahinifarahani, Terry Yue Zhuo, Shweta Garg, Baishakhi Ray, Rajdeep Mukherjee, Varun Kumar. arXiv 2026.
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All — Chenxi Huang, Alex Mathai, Feiyang Yu, Aleksandr Nogikh, Petros Maniatis, Franjo Ivancic, Eugene Wu, Kostis Kaffes, Junfeng Yang, Baishakhi Ray. ICML 2026.
Quantifying Trust: Financial Risk Management for Trustworthy AI Agents — Wenyue Hua, Tianyi Peng, Chi Wang, Jiaxin Pei, Ian Kaufman, Bryan Lim, Chandler Fang. arXiv 2026.
Estimating Tail Risks in Language Model Output Distributions — Rico Angell, Raghav Singhal, Zachary Horvitz, Zhou Yu, Rajesh Ranganath, Kathleen McKeown, He He. ICML 2026.
Conversational Customization of Productivity Systems: A Design Probe of Malleable AI Interfaces — Karthik Sreedhar, Aryan Kaul, Lydia Chilton. arXiv 2026.
AutoRPA: Efficient GUI Automation through LLM-Driven Code Synthesis from Interactions — Minghao Chen, Xinyi Hu, Zhou Yu, Yufei Yin. ICML 2026.
Agents for Web Testing: A Case Study in the Wild — Naimeng Ye, Xiao Yu, Ruize Xu, Tianyi Peng, Zhou Yu. LAw Workshop at NeurIPS 2025.
Law-Abiding Agents: Demonstrating Safe Tax Preparation with Data Flow Control — Charlie Summers, Prajwal Raghunath, Zhibin Shen, Hardik Gupta, Eric Choi, Mayur Kulkarni, Sally Go, Peter Yu, Akriti Agarwal, Zhuo Zhang, Oliver Kennedy, Eugene Wu.*Currently in preprint 2026.
VISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation — Yunan Lu, Ryan Shea, Yusen Zhang, Zhou Yu. arXiv 2026.
PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures — Yunan Lu, Luigi Liu, Omar Yahia, Arpit Sharma, Zhou Yu. arXiv 2026.