Datasets:
Request access to PersonalizationV4
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
PersonalizationV4 is an evaluation benchmark of fictional users. Access is gated to keep its test questions out of web-scale training corpora. Please use the dataset for research and evaluation, do not repost its questions in plain text on the web, do not use it to profile or identify real people, and cite it in work that uses it.
Log in or Sign Up to review the conditions and access this dataset content.
PersonalizationV4
PersonalizationV4 (PV4) is a synthetic personalization benchmark. Each user is a detailed fictional persona who has had 200 short conversations with an AI assistant. The evaluation questions place the user in a new scenario and ask what they would most likely do or prefer, and each one is written to require combining at least two facts about the user. A model never sees the persona itself: it gets the user's conversations, in which those traits are shown rather than stated, as its memory and answers in free text.
PV4 is part of Memorilla, where the generation pipeline (pv4/) and the
evaluation code live.
| Users | 149 (119 training, 30 evaluation; disjoint) |
| Conversations per user | 200 two-turn chats (one user message, one assistant reply) |
| Questions per user | 133-145 (10 categories per persona) |
| Rows | train 15,058, validation 1,755, test 4,237 |
| Hard test subset | 629 questions |
| Scoring | nearest of five candidate answers under Qwen3-Embedding-4B |
Usage
Accept the access terms on this page, authenticate with hf auth login, then:
from datasets import load_dataset
pv4 = load_dataset("MemoryAsModality/PersonalizationV4")
row = pv4["test"][0]
print(row["question"])
print(row["answer"])
print(row["documents"][0])
To evaluate a Memorilla checkpoint or a retrieval baseline, use the main repository:
python evaluate.py --benchmark pv4 --checkpoint runs/personalization_pv4/epoch-04
python evaluate_baselines.py --benchmark pv4 --method rag --top_k 5
Dataset structure
| column | type | description |
|---|---|---|
user_id |
int64 | User the question is about; all rows of a user share the same documents. |
question |
string | Scenario-based question about the user. |
answer |
string | Reference answer (one of choices). |
choices |
list[string] | The five candidate answers in A-E order; a missing candidate is an empty string and is skipped by the scorer. |
documents |
list[string] | The user's 200 chats in topic order, each a Leo: <user message> turn followed by an Assistant: <reply> turn (Leo is a fixed speaker tag for the user). |
hard |
bool | True for questions in the hard subset (test split only). |
| split | users | rows | content |
|---|---|---|---|
train |
119 | 15,058 | questions of the training users, minus a 10% per-user held-out part |
validation |
119 (same as train) | 1,755 | the held-out 10% of each training user's questions |
test |
30 (unseen) | 4,237 | every question of the evaluation users |
The per-user hold-out draws a permutation of the user's questions with numpy.random.default_rng(23) and holds out the
first ceil(0.1 * n); this is the rule of datasets.Dataset.train_test_split(test_size=0.1, seed=23). Training users
have ids between 6 and 125 and evaluation users ids 126 to 155.
Scoring
Every question comes with five candidate answers: the reference answer and four distractors. The distractors are designed to be equally plausible choices for a reasonable person, and the five candidates are matched in length, grammatical structure and specificity, so the reference cannot be singled out from the candidates alone.
The model never sees the candidates. It reads the question (with the user's conversations available as memory) and generates a free-text answer. The generation and the non-empty candidates are embedded with Qwen3-Embedding-4B; the answer is correct when the candidate with the highest cosine similarity to the generation is the reference. Accuracy is the mean over questions.
Hard subset. The hard column marks 629 test questions that the untrained Qwen3-4B-Instruct-2507 decoder gets
wrong in at least one of two settings: with no documents, or with only the single most relevant conversation
retrieved. Accuracy on these questions is reported as accuracy_hard.
Generation
Each user starts from a long-form persona profile (about 2,400 words on average) that expands a five-sentence seed persona from Synthetic-Persona-Chat into a detailed life: identity, work, family and friends, hobbies and tastes, personality, daily routine and a secret project.
| Step | Output (per user) | Model | Reasoning effort |
|---|---|---|---|
| 1. Question categories: the 10 most testable dimensions of the persona | categories.txt |
gpt-5.1 |
none |
| 2. Chat topics: 200 one-sentence scenarios covering every facet of the persona, early skeptical and later reliant phases, and requests secretly related to hidden projects | chat_topics.txt |
gpt-5.1 |
none |
| 3. Chats: one two-turn chat per topic; the user's traits are shown, never stated | chats/{topic}.txt |
gpt-5-mini |
minimal |
| 4. Questions: 15 requested per category, each with a reference answer, four distractors and a rationale | qa/{category}.txt |
gpt-5.1 |
none |
5. Question table: parse the questions, drop malformed ones, and move each reference to a random letter with random.Random(42) |
qa.csv |
||
| 6. Dataset: per-user splits, chats attached as documents, hard-subset flags | data/*.parquet |
The question prompt asks for scenario-embedded questions that require combining two or more persona facts, distractors
that a reasonable person might genuinely prefer (including the best practice this persona rejects), and candidates
matched in length, structure and specificity. The exact prompts and the code of every step are in the pv4/ directory
of the Memorilla repository.
Raw files
raw/ holds every intermediate output of the pipeline for the 149 released users:
raw/
personas/user_N.txt persona profile of user N (pipeline input)
persona_seeds.csv user_id, seed (Synthetic-Persona-Chat persona)
hard_subset.csv user_id, question_index of the 629 hard test questions
users/user_N/
categories.txt question categories, one per line (line i <-> qa/i.txt)
chat_topics.txt chat topics, one per line (line i <-> chats/i.txt)
chats/{i}.txt two-turn chat for topic i
qa/{i}.txt generated questions for category i, as returned by the model
qa.csv parsed questions: index, category, question, choice_a..choice_e, correct_choice, rationale
In qa.csv, correct_choice is the letter after the answer
shuffle, while rationale is the model's explanation and refers to the letters in qa/{i}.txt.
data/ is rebuilt from raw/ with:
python -m pv4.build_dataset --users_dir raw/users --hard_subset raw/hard_subset.csv --output_dir data
License and attribution
- PV4 is released under CC BY 4.0.
- The seed personas come from Google's Synthetic-Persona-Chat (Jandaghi et al., 2023), released under CC BY 4.0.
- All text in PV4 (persona profiles, chat topics, chats, questions and answers) was generated with OpenAI models.
- All users are fictional; any resemblance to real people is coincidental.
Citation
@misc{memorilla2026,
title = {Memorilla: Memory as a Modality for LLMs},
author = {Memorilla Team},
year = {2026},
url = {https://github.com/snap-stanford/memorilla}
}
@article{jandaghi2023faithful,
title = {Faithful Persona-based Conversational Dataset Generation with Large Language Models},
author = {Jandaghi, Pegah and Sheng, XiangHai and Bai, Xinyi and Pujara, Jay and Sidahmed, Hakim},
journal = {arXiv preprint arXiv:2312.10007},
year = {2023}
}
- Downloads last month
- 262