Back to Feed
Agents / Benchmarks & Evals

Improving Tool Use With Looped Language Models

Original: Looped Language Models Improve Compositional Tool Calling

Listen to the summary

Uses a voice available on your device

Audio options
On this page 4 sections
Related concepts 7 concepts

Key Takeaways

  • Ouro-1.4B and Ouro-2.6B models demonstrate significant improvements in tool-calling benchmarks including BFCL.
  • The models excel particularly in Parallel and Parallel-Multiple categories within the BFCL benchmark.
  • Increasing recurrent depth enhances performance in nested workflows, with Ouro-2.6B reaching a 0.371 win rate on NESTful.
  • Retrofitted recurrent models perform worse than natively recurrent models when handling deeply nested tool-use workflows.

Summary & Methodology Analysis

The researchers introduced Ouro-1.4B and Ouro-2.6B, models that refine latent representations through the recurrent application of a shared stack of Transformer layers. This approach enables iterative computation during test time rather than relying on a fixed depth. The 1.4B version utilizes a 24-layer recurrent stack, while the 2.6B version doubles this to 48 layers through continued pretraining. All models were evaluated using supervised fine-tuning, a process where models are trained on specific input-output examples, using the Hermes function calling dataset to align their output with API execution requirements. This architecture aims to solve compositional tool-calling challenges where models must manage state across multiple dependencies.

Interactive System Flowchart

Click diagram to expand and zoom

Cross-Examination & FAQs

A deeper dive clarifying mechanics, constraints, and baseline evaluations.

Q1. What is the main goal of this research?

The research aims to improve how language models handle compositional tool-calling, which involves coordinating multiple API calls and maintaining state across interactions.

Q2. What are Ouro models?

Ouro models are language models that employ iterative latent computation through a shared stack of Transformer layers to refine their decision-making during test time.

Q3. How were the models trained?

The models underwent supervised fine-tuning using the Hermes function calling dataset.

Q4. What benchmarks were used to evaluate the models?

The researchers used BFCL v3, NESTful, and API-Bank to assess model performance across various tool-use tasks.

Q5. How does recurrent depth affect performance?

Increasing recurrent depth generally improves performance on hierarchical tool use, specifically in NESTful workflows.

Q6. How do retrofitted models compare to natively recurrent models?

Retrofitted models are substantially weaker than natively recurrent models when handling deeply nested tool-use workflows.

Q7. What are the limitations of the current study?

The study is restricted to static, single-turn evaluation benchmarks and does not currently address live, multi-turn agentic environments.

Q8. What specific metrics did Ouro-2.6B achieve on the BFCL benchmark?

Following supervised fine-tuning, Ouro-2.6B achieved 86.4% performance on the overall BFCL benchmark.

Q9. Does the paper specify the inference latency of these models?

The paper does not specify the inference latency of the Ouro models.

Flag an issue

What is wrong with this summary?

What is wrong?