Model Releases arXiv cs.AI

Pistis Technical Report

Pistismultimodal LLMIDRLQwen

The Pistis model family comprises 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. That framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT).

Building on the SFT foundation, the authors propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives—rather than optimizing either in isolation or combining them in a static joint loss—IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long-horizon agentic trajectories.

At both model scales, the framework produces two specialized variants: Pistis-Thinking, designed to strengthen deep multimodal reasoning, and Pistis-Agentic, which additionally incorporates agentic trajectory data to support long-horizon planning, iterative reasoning, and tool use. Pistis-Agentic is particularly strong in multimodal search, and both scales outperform their corresponding base models.

Beyond model-parameter optimization, the authors introduce Pistis-Auto-Harnessing (PAH), a system-level method that automatically improves the agent's inference harness through iterative optimization. Experiments demonstrate that PAH enhances model performance without updating the model parameters or increasing the interaction budget.

Read original →

← Back to home