Request access to PersonalizationV4

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

PersonalizationV4 is an evaluation benchmark of fictional users. Access is gated to keep its test questions out of web-scale training corpora. Please use the dataset for research and evaluation, do not repost its questions in plain text on the web, do not use it to profile or identify real people, and cite it in work that uses it.

Log in or Sign Up to review the conditions and access this dataset content.

PersonalizationV4

PersonalizationV4 (PV4) is a synthetic personalization benchmark. Each user is a detailed fictional persona who has had 200 short conversations with an AI assistant. The evaluation questions place the user in a new scenario and ask what they would most likely do or prefer, and each one is written to require combining at least two facts about the user. A model never sees the persona itself: it gets the user's conversations, in which those traits are shown rather than stated, as its memory and answers in free text.

PV4 is part of Memorilla, where the generation pipeline (pv4/) and the evaluation code live.

Users 149 (119 training, 30 evaluation; disjoint)
Conversations per user 200 two-turn chats (one user message, one assistant reply)
Questions per user 133-145 (10 categories per persona)
Rows train 15,058, validation 1,755, test 4,237
Hard test subset 629 questions
Scoring nearest of five candidate answers under Qwen3-Embedding-4B

Usage

Accept the access terms on this page, authenticate with hf auth login, then:

from datasets import load_dataset

pv4 = load_dataset("MemoryAsModality/PersonalizationV4")
row = pv4["test"][0]
print(row["question"])
print(row["answer"])
print(row["documents"][0])

To evaluate a Memorilla checkpoint or a retrieval baseline, use the main repository:

python evaluate.py --benchmark pv4 --checkpoint runs/personalization_pv4/epoch-04
python evaluate_baselines.py --benchmark pv4 --method rag --top_k 5

Dataset structure

column type description
user_id int64 User the question is about; all rows of a user share the same documents.
question string Scenario-based question about the user.
answer string Reference answer (one of choices).
choices list[string] The five candidate answers in A-E order; a missing candidate is an empty string and is skipped by the scorer.
documents list[string] The user's 200 chats in topic order, each a Leo: <user message> turn followed by an Assistant: <reply> turn (Leo is a fixed speaker tag for the user).
hard bool True for questions in the hard subset (test split only).
split users rows content
train 119 15,058 questions of the training users, minus a 10% per-user held-out part
validation 119 (same as train) 1,755 the held-out 10% of each training user's questions
test 30 (unseen) 4,237 every question of the evaluation users

The per-user hold-out draws a permutation of the user's questions with numpy.random.default_rng(23) and holds out the first ceil(0.1 * n); this is the rule of datasets.Dataset.train_test_split(test_size=0.1, seed=23). Training users have ids between 6 and 125 and evaluation users ids 126 to 155.

Scoring

Every question comes with five candidate answers: the reference answer and four distractors. The distractors are designed to be equally plausible choices for a reasonable person, and the five candidates are matched in length, grammatical structure and specificity, so the reference cannot be singled out from the candidates alone.

The model never sees the candidates. It reads the question (with the user's conversations available as memory) and generates a free-text answer. The generation and the non-empty candidates are embedded with Qwen3-Embedding-4B; the answer is correct when the candidate with the highest cosine similarity to the generation is the reference. Accuracy is the mean over questions.

Hard subset. The hard column marks 629 test questions that the untrained Qwen3-4B-Instruct-2507 decoder gets wrong in at least one of two settings: with no documents, or with only the single most relevant conversation retrieved. Accuracy on these questions is reported as accuracy_hard.

Generation

Each user starts from a long-form persona profile (about 2,400 words on average) that expands a five-sentence seed persona from Synthetic-Persona-Chat into a detailed life: identity, work, family and friends, hobbies and tastes, personality, daily routine and a secret project.

Step Output (per user) Model Reasoning effort
1. Question categories: the 10 most testable dimensions of the persona categories.txt gpt-5.1 none
2. Chat topics: 200 one-sentence scenarios covering every facet of the persona, early skeptical and later reliant phases, and requests secretly related to hidden projects chat_topics.txt gpt-5.1 none
3. Chats: one two-turn chat per topic; the user's traits are shown, never stated chats/{topic}.txt gpt-5-mini minimal
4. Questions: 15 requested per category, each with a reference answer, four distractors and a rationale qa/{category}.txt gpt-5.1 none
5. Question table: parse the questions, drop malformed ones, and move each reference to a random letter with random.Random(42) qa.csv
6. Dataset: per-user splits, chats attached as documents, hard-subset flags data/*.parquet

The question prompt asks for scenario-embedded questions that require combining two or more persona facts, distractors that a reasonable person might genuinely prefer (including the best practice this persona rejects), and candidates matched in length, structure and specificity. The exact prompts and the code of every step are in the pv4/ directory of the Memorilla repository.

Raw files

raw/ holds every intermediate output of the pipeline for the 149 released users:

raw/
  personas/user_N.txt            persona profile of user N (pipeline input)
  persona_seeds.csv              user_id, seed (Synthetic-Persona-Chat persona)
  hard_subset.csv                user_id, question_index of the 629 hard test questions
  users/user_N/
    categories.txt               question categories, one per line (line i <-> qa/i.txt)
    chat_topics.txt              chat topics, one per line (line i <-> chats/i.txt)
    chats/{i}.txt                two-turn chat for topic i
    qa/{i}.txt                   generated questions for category i, as returned by the model
    qa.csv                       parsed questions: index, category, question, choice_a..choice_e, correct_choice, rationale

In qa.csv, correct_choice is the letter after the answer shuffle, while rationale is the model's explanation and refers to the letters in qa/{i}.txt.

data/ is rebuilt from raw/ with:

python -m pv4.build_dataset --users_dir raw/users --hard_subset raw/hard_subset.csv --output_dir data

License and attribution

  • PV4 is released under CC BY 4.0.
  • The seed personas come from Google's Synthetic-Persona-Chat (Jandaghi et al., 2023), released under CC BY 4.0.
  • All text in PV4 (persona profiles, chat topics, chats, questions and answers) was generated with OpenAI models.
  • All users are fictional; any resemblance to real people is coincidental.

Citation

@misc{memorilla2026,
  title  = {Memorilla: Memory as a Modality for LLMs},
  author = {Memorilla Team},
  year   = {2026},
  url    = {https://github.com/snap-stanford/memorilla}
}

@article{jandaghi2023faithful,
  title   = {Faithful Persona-based Conversational Dataset Generation with Large Language Models},
  author  = {Jandaghi, Pegah and Sheng, XiangHai and Bai, Xinyi and Pujara, Jay and Sidahmed, Hakim},
  journal = {arXiv preprint arXiv:2312.10007},
  year    = {2023}
}
Downloads last month
262

Paper for MemoryAsModality/PersonalizationV4