Research arXiv cs.AI

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

business intelligenceLLM agentsbenchmarkpost-training

Business intelligence is a cornerstone of enterprise decision-making and widely used in tools like Power BI and Tableau. In traditional BI workflows, users must first identify relevant tables, perform data transformations, and build join relationships before they can answer their business questions. Those preparation steps are complex and time-consuming, which makes BI challenging.

Given LLMs' strong capabilities in working with data, this work studies whether they can answer BI questions end-to-end without users manually performing the tedious preparation steps. The authors harvest a large collection of real-world BI projects from public sources and manually extract pairs of (questions, ground-truth answers) from real user dashboards. The resulting benchmark, BI-Bench, is described as the first benchmark to systematically study LLMs' ability on end-to-end BI.

On BI-Bench, even frontier LLMs perform poorly, with less than 50% accuracy. To address these limitations, the authors design a tool-augmented BI-Agent that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages. They also develop a post-training framework that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL).

BI-Agent achieves substantial accuracy gains of up to 40 percentage points with vanilla LLMs, while the post-trained BI-Agent yields gains of up to 30 points. The results highlight the importance of combining tool-augmented reasoning with domain-specific post-training in complex BI workflows and point to promising directions for future research.

Read original →

← Back to home