Research arXiv cs.LG

Reinforcement Learning of Communication in a Mesh of Small Language Models

multi-agentsmall language modelsreinforcement learningconfidence scoring

Language models can gain accuracy from more test-time compute, but majority voting over independent samples saturates: as the number of samples grows, the vote converges to the model’s most frequent answer. Communication can add something sampling cannot—an agent that solves a problem can pass the key step to the others.

TalkMesh is a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head; the most confident agent broadcasts a hint, and agents below a confidence threshold revise, keeping each revision that outscores its proposal. Gossip consensus approximates the vote weighted by confidence without a coordinator. A talk policy, trained with group relative policy optimization on the change in correctness after revision, writes the hints and revisions.

With three agents, which together generate at most six outputs, the mesh reaches the accuracy of majority voting over 32 samples with each of three models. Trained with at most 8 agents and evaluated with 32, it raises accuracy from 0.568 under self-consistency to 0.705 (Qwen3.5-0.8B, GSM8K) and from 0.492 to 0.722 (SmolLM3-3B, MATH-500).

In an adversarial setting, when 4 of 8 agents collude on a wrong answer with fabricated confidence and poisoned hints, majority vote accuracy falls to 0.000 (Qwen3.5-0.8B, GSM8K). A defended mesh, whose agents rescore solutions with their own confidence heads, retains 0.507. Across reasoning, embodied coordination, and traffic signal control, the results indicate that messages improve a decision when the acting agent cannot observe the information it requires and another agent can send it.

Read original →

← Back to home