Skip to content
TrackPodcasts
scienceJan 15, 202622:06pending

KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions

About this episode

🤗 Upvotes: 47 | cs.AI, cs.IR

Authors:
Tingyu Wu, Zhisheng Chen, Ziyan Weng, Shuhe Wang, Chenglong Li, Shuo Zhang, Sen Hu, Silin Wu, Qizhen Lan, Huacan Wang, Ronghao Chen

Title:
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions

Arxiv:
http://arxiv.org/abs/2601.04745v1

Abstract:
Existing long-horizon memory benchmarks mostly use multi-turn dialogues or synthetic user histories, which makes retrieval performance an imperfect proxy for person understanding. We present \BenchName, a publicly releasable benchmark built from long-form autobiographical narratives, where actions, context, and inner thoughts provide dense evidence for inferring stable motivations and decision principles. \BenchName~reconstructs each narrative into a flashback-aware, time-anchored stream and evaluates models with evidence-linked questions spanning factual recall, subjective state attribution, and principle-level reasoning. Across diverse narrative sources, retrieval-augmented systems mainly improve factual accuracy, while errors persist on temporally grounded explanations and higher-level inferences, highlighting the need for memory mechanisms beyond retrieval. Our data is in \href{KnowMeBench}{https://github.com/QuantaAlpha/KnowMeBench}.

Get every episode summarized

Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions

Daily Paper Cast

0:00
22:06

More episodes

More from Daily Paper Cast

View all episodes →