Researchers introduce DORA Explorer, a training-free, inference-time algorithm that improves exploration in large language model agents. It generates multiple candidate actions, scores them using sequence-level log probabilities, and samples actions via an exploration parameter. The work addresses LLM agents’ difficulty producing diverse, non-repetitive action sequences in sequential decision-making tasks. Existing approaches, such as temperature-based sampling, operate at the token level and often yield insufficient exploration and looping behaviors. Evaluations show DORA substantially outperforms temperature sampling in Multi-Armed Bandit experiments and improves TextWorld performance in the TALES benchmark, notably boosting Qwen-2.5 7B’s ReAct score from 31.43% to 45.5%. Code and demos are publicly available.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.