DAGO learns which workflows to fuse
The expensive question hiding inside automated design
Agentic workflows are programs that coordinate large language model calls, external tools, and control-flow logic to solve tasks that no single model invocation handles well. When researchers automate the design of these workflows, they convert construction into a search problem over executable programs. That conversion sounds clean. It is not, because every candidate the search proposes must be run — its LLM calls executed, its tools invoked, its outputs scored against validation examples — before anyone knows whether it was worth proposing.
DAGO, short for Directed Acyclic Graph Optimization, begins from this cost structure. DAGO treats the evaluation budget as the scarce resource and asks a narrower question than most workflow-search papers: given that fusing two or more previously discovered workflows can reuse good design, how do you decide which combinations deserve the next expensive test? The answer it develops is a contextual-bandit framework in which each candidate parent combination is an arm, the reward is the validation score of the child that fusion produces, and a diagonal LinUCB policy learns across arms which combinations tend to yield strong offspring.
The gain is concrete. Across six benchmarks covering mathematical reasoning, code generation, and question answering, DAGO recorded the highest macro-average score among the evaluated baselines. [1] Under the same validation-evaluation budget cap that AFlow operates under, it lifted the macro-average from 80.3 to 81.7 while cutting aggregate search expenditure by 11.2 percent. [1] That is a small number and a real one: better results for less compute, on the same budget.
Why an effective fusion operator is not enough
Multi-parent fusion already exists in the literature, including EvoFlow and MermaidFlow. EvoFlow uses LLM-based crossover, and MermaidFlow introduces constraint-preserving crossover over workflow graphs. [1] Both establish that recombining pieces of previously discovered workflows can produce something better than either parent alone. [1] A workflow with strong problem decomposition can supply guidance to a workflow with a strong verification procedure, and the child may inherit the useful parts of both.
What those operators do not settle is the selection problem sitting one level up. High individual validation scores do not guarantee that a given pair will produce a strong child, and the number of possible combinations grows combinatorially as the archive of discovered workflows expands. Evaluating every combination is impractical. Choosing combinations at random throws away the signal from past fusion outcomes. The central challenge DAGO names is therefore to learn which combinations merit the next expensive fusion and evaluation, under a budget that makes exhaustive testing impossible.
The mechanism that turns a score into a decision

DAGO’s design follows from taking that challenge literally. Each workflow is represented by pretrained embeddings of its code and prompts. An arm — a candidate combination of parents — is represented by concatenating its parents’ embeddings in descending order of validation score, with node identifiers breaking ties. That canonical ordering removes arbitrary permutations of the same parent set, so the full arm space contains a binomial number of choices before each round. [1] Rather than learning an independent reward estimate for every possible combination, DAGO shares one linear reward model across arms. A diagonal LinUCB policy scores each proposed arm by combining predicted child-workflow quality with an uncertainty-driven exploration bonus. The shared representation lets feedback from arms already evaluated inform the ranking of newly proposed arms, while the exploration term pushes the search toward alternatives beyond those with the highest current predictions. The diagonal approximation keeps scoring and online updates lightweight.
Candidate arms are not enumerated exhaustively. At each iteration, annealed score-based sampling proposes a manageable set, using parent scores to guide which combinations get considered. Then the bandit selects among them. The division of labor matters: parent scores shape the proposal, and observed fusion rewards shape the selection. Evaluation opportunities are allocated through an explicit exploration-exploitation mechanism rather than through a fixed rule.
What happens after the arm is pulled
Pulling an arm generates one child. An LLM summarizes the selected parents’ strengths and constraints, then performs summary-guided fusion, starting from the strongest parent and applying targeted modifications informed by the others. [1] The child is validated, and its validation score becomes the reward that updates the bandit. [1] The child is then added to a shared directed acyclic graph that stores discovered workflows and their multi-parent generation lineage.
That graph is the search history, not the execution graph inside any individual workflow. Each edge records that a parent was supplied when generating a child; it does not certify that a specific component was inherited. Nodes store code, prompts, scores, and cached embeddings. As the graph grows, so does the pool of parents available for subsequent arm proposals. The loop is closed: evaluation produces reward, reward improves selection, selection produces new workflows, and new workflows expand the set of combinations worth considering. At the end of search, DAGO returns the workflow with the highest validation score for reuse on the target task. Arm selection happens during optimization, not separately for each test query.
The numbers behind the claim
The six benchmarks span mathematical reasoning on GSM8K and MATH, code generation on HumanEval and MBPP, and question answering on HotpotQA and DROP. DAGO’s macro-average improvement over AFlow, from 80.3 to 81.7, was achieved under matched validation-evaluation budgets, and the 11.2 percent reduction in aggregate search expenditure was measured across the same six benchmarks.
Ablation studies isolate the contribution of the learned selection. On all three tested benchmarks, the default LinUCB policy achieved higher mean scores than random selection and higher mean scores than an exploration-free variant. [1] That second comparison is the more informative one: removing the exploration bonus degrades performance, which supports the claim that the balance between exploiting high-predicted arms and exploring uncertain ones is doing real work, not decorating the method.

What the paper does not claim
DAGO does not assume that the expected reward of an arm is linear in its features, and it does not assume that reward distribution is stationary across rounds. The shared linear model is a working approximation, chosen to make learning from limited feedback feasible, not a statement about how fusion quality actually behaves. The paper is explicit that a stored score is an observed evaluation, not the unknown expected performance of that workflow. Validation feedback is available during search; test data are not used for generation, arm selection, or stopping.
These are boundaries, and they matter for reading the result correctly. The 1.4-point macro-average gain is a gain under a specific budget constraint, against specific baselines, on six benchmarks. The 11.2 percent expenditure reduction is measured under matched budgets. Neither number says that DAGO finds the best possible workflow, only that it allocates a limited evaluation budget better than the alternatives it was compared against.
The question underneath the improvement
There is a quiet assumption running through the whole design: that the workflow with the highest validation score is the one worth returning. DAGO optimizes toward arg max over the discovered set, and the paper’s objective is stated in exactly those terms — discover a high-scoring workflow within the evaluation budget. The machinery appears built to serve that target efficiently.
But the search is also building something else along the way: a directed acyclic graph of lineage, in which every child carries a record of which parents were fused to produce it. The paper notes that edges record supply, not inheritance. A parent can be listed without any of its components surviving into the child. So the graph is a history of decisions, not a map of what actually worked. In the paper’s own framing, the objective is to discover a high-scoring workflow within the evaluation budget, and per-round reward prediction is a search surrogate rather than an equivalent best-workflow guarantee. [1] Whether that distinction matters for how we evaluate automated design — whether the goal should be the single best artifact or an accurate account of why it is best — is not a question the benchmarks can answer. DAGO makes the selection problem explicit and solves it well. It leaves the target itself exactly where it found it.
