Alibaba AI Chief Says Months of Data Work Can Be Done with Natural-Language Instructions; Agentic AI Takes Center Stage in Jeju
- Input
- 2026-08-12 10:15:04
- Updated
- 2026-08-12 10:15:04

[Financial News, Jeju = Reporter Jung Yong-bok] The era is opening in which data engineering tasks that once took months can now be defined and automated through natural-language instructions. Moving away from the old approach of having people manually find and process data before asking AI questions, Jeju has emerged as a new focal point for the "agentic data stack," where AI agents handle data collection, cleansing, querying and analysis in sequence.
Jingren Zhou, Senior Vice President and Chief AI Architect at Alibaba Group, delivered a keynote speech at ACM KDD 2026 in Jeju on the morning of the 12th. His talk was titled "The Agentic Data Stack: How LLMs Power Data Engineering and Orchestration."
The key point Zhou presented was that large language models (LLMs) and AI agents are no longer just tools that assist data analysis. They are changing the data processing workflow itself.
Until now, companies and research institutions have had to build data pipelines tailored to specific tasks and repeatedly rely on specialists to find, organize and transform the necessary information. Zhou explained, "By applying LLMs and AI agents to core tasks such as data transformation, schema discovery, text-to-SQL and feature engineering, these processes can be automated through natural-language instructions."
He also suggested that data engineering work that once took months can be greatly reduced by having people describe the desired conditions and outcomes in natural language. Instead of designing every procedure and program manually, users can simply ask, "Analyze this data in this way," and AI will assemble and execute the required steps.
The technical terms become clearer when broken down. "Schema discovery" refers to identifying what tables and fields exist in a large database and how they are connected. "Text-to-SQL" is a technology that converts everyday language questions into SQL commands that a database can understand and use to retrieve the needed information.
"Feature engineering" is the process of extracting meaningful variables and attributes from raw data so that AI or statistical models can use them for training and analysis. In other words, LLMs are expanding into areas that once depended heavily on the experience and repetitive work of data specialists.
Large-scale data preprocessing required for training AI foundation models was also highlighted as a major challenge. As AI increasingly handles not only text but also images, video and audio, the scale and complexity of multimodal data that must be processed before training are growing rapidly.
Zhou explained, "By combining LLMs' ability to understand language and meaning with existing database technologies, we can build pipelines that process large-scale multimodal data more efficiently and flexibly while also reflecting the meaning contained in the data."
He described "orchestration" as the next stage in data processing. Rather than a structure in which a single AI model simply receives a question and returns an answer, AI selects the necessary tools, sets the order of tasks and checks intermediate results while linking multiple steps together.
Zhou said, "When LLMs are combined within an agent framework with context, memory, tools and verification mechanisms, long-horizon reasoning and execution across multiple steps become possible."
As a real-world example, he introduced Alibaba Group's AI agent framework, AgentScope. Data agents built on AgentScope are designed to autonomously collect the required data, select and refine it, query databases and analyze the results. He said the system can be used to support business intelligence (BI) and decision-making in companies.
In simple terms, the old model was "humans prepare the data and then ask AI questions." In the agentic data stack, AI's role becomes much broader. The agent connects the entire process, from finding where the needed data is stored to deciding how it should be processed and producing the final analysis.
Zhou is the person leading the development of Alibaba Group's core AI technologies. As Senior Vice President and Chief AI Architect, he created major AI foundation models, including the Qwen and Wan series, and now oversees their development and application across Alibaba Group's businesses.
Before that, he served as Chief Technology Officer at Alibaba Cloud, where he oversaw cloud computing technology and product development. He also took part in developing AI technologies and infrastructure for personalized search, product recommendations and advertising on Alibaba Group's e-commerce platforms.
Before joining Alibaba Group, he worked on big data and database research and development at Microsoft. He earned a doctorate in computer science from Columbia University and is a Fellow of both the Association for Computing Machinery (ACM) and IEEE.
The lecture showed that the focus of AI competition is expanding beyond model performance itself to how effectively companies and institutions can let AI independently find, process and use the vast amounts of data they hold.
In particular, if the trend of redesigning human-built data pipelines around natural-language instructions and AI agents spreads further, the role of data specialists is likely to shift away from repetitive processing and toward data quality verification, workflow design and outcome assessment.
ACM KDD 2026, a leading international conference in data science and AI, has been held at the International Convention Center Jeju (ICC Jeju) from the 9th to the 13th.
[email protected] Reporter Jung Yong-bok Reporter