LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
Introduced in arXiv paper 2608.05246, LUNAR is described as the first benchmark for cross-domain behavioral personalization based on app interaction histories. Unlike prior approaches that use small behavioral signals, LUNAR grounds evaluation in heterogeneous daily-life activities, covering universal domains like communication, shopping, and entertainment. The benchmark aims to test whether LLMs can infer and adapt to user preferences from long-term behavior logs, a more realistic personalization scenario. Its release could push model developers toward better user modeling and context-aware response generation. The paper likely includes baseline results and exposes current LLM shortcomings in behavioral personalization.