hello, world

I'm Quan Liu.

LLM & Agent Research Scientist · Tech Lead @ Accenture

I'm a Sr. Research Scientist and Tech Lead at Accenture's Center for Advanced AI, and a founding member of its agentic-systems team. I lead the team building our coding agent and automated agent-generation pipeline, and shipped core features of AIRefinery — Accenture's enterprise multi-agent platform, featured at NVIDIA GTC. My research on tool-using agents, retrieval, and post-training alignment appears at ICLR and NeurIPS; my earlier vision and medical-imaging work was published at CVPR, MICCAI, and in Nature Nanotechnology.

About

Hi, I'm Quan.

I'm an LLM and agent researcher, and an engineer at heart. I like turning open research questions into systems that actually run — and I care as much about whether something holds up in production as whether it works in a paper.

Today I'm at Accenture's Center for Advanced AI, where, as a founding member of our agentic systems team, I lead a team building a coding agent and an automated agent-generation pipeline, and contribute core features to AIRefinery. I earned my Ph.D. in Computer Science at Vanderbilt, working across medical imaging, computational vision, and machine learning.

My research spans LLM agents, tool use, retrieval, and post-training alignment, with recent work at ICLR and NeurIPS. Off the clock, I'm happiest on a basketball court or somewhere along the California coast.

Career

Experience

May 2024 — Present

Sr. Research Scientist / Tech Lead

Accenture — Center for Advanced AI · Mountain View, CA
  • Founding member of the agentic-systems team; lead the team building a coding agent and an automated agent-generation pipeline.
  • Shipped core features of AIRefinery, Accenture's enterprise multi-agent platform (featured at NVIDIA GTC).
  • Built schema-constrained tool execution (MCP-style grounding, validation, and tracing) and a hybrid retrieval stack (E5 + BM25 + FAISS) powering production agents.
  • Ran post-training alignment — SFT, PPO, and GRPO — improving reasoning, tool planning, and multi-agent coordination.
May – Aug 2023

Research Intern

Merck — Image Data Analytics · West Point, PA
May – Aug 2022

Research Intern

Alibaba — DAMO Academy · New York, NY
Background

Education

Aug 2020 — May 2024

Ph.D., Computer Science

Vanderbilt University · Nashville, TN
Aug 2018 — May 2020

Ph.D. study, Computer Engineering

Case Western Reserve University · Cleveland, OH
Sep 2014 — Jun 2018

B.S., Electrical Engineering

Huazhong University of Science and Technology · Wuhan, China
Selected research

Recent work & papers

All on Google Scholar
2026
ICLR 2026 · NeurIPS 2025 Workshop

MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Zhenting Wang, Qi Chang, Hemani Patel, Shashank Biju, Cheng-En Wu, Quan Liu, Aolin Ding, Alireza Rezazadeh, Ankit Shah, Yujia Bao, Eugene Siow
2026
ICLR 2026

DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning

Yaxuan Wang, Chris Yuhao Liu, Quan Liu, Jinglong Pang, Wei Wei, Yujia Bao, Yang Liu
2025
arXiv preprint

PromptBridge: Cross-Model Prompt Transfer for Large Language Models

Yaxuan Wang, Quan Liu, Zhenting Wang, Zichao Li, Wei Wei, Yang Liu, Yujia Bao
Things I build

Featured projects

All projects
Lead project

Coding agent & agent auto-generation

As a founding member, I lead a team building a coding agent and an automated pipeline that generates task-specific agents — taking a spec to a working, tool-grounded agent with minimal manual wiring.

Enterprise platform

AIRefinery

Accenture's enterprise multi-agent platform, also available as an open-source Python SDK. I contributed core feature implementation across the agent stack. Featured at NVIDIA GTC.

Benchmark

MCP-Bench

An open benchmark for tool-using LLM agents on complex, multi-step tasks over live MCP servers. ICLR 2026 / NeurIPS 2025 Workshop.

What's new

Latest updates

All updates
Let's work together

Open to collaboration

I'm always glad to talk with people building reliable LLM agents — whether that's a research collaboration, an open benchmark, or just comparing notes on what actually holds up in production.

How we can collaborate