The Dataset Viewer has been disabled on this dataset.

Astro Sky Image VQA

A visual-question-answering benchmark of synthetic astronomical figures with exact, machine-generated ground truth.

Every figure here was drawn by us with matplotlib, so for each one we know the things a reader of a paper can only estimate: the true minimum of the colour axis, the real angular width of the field, the exact axis limits, the colormap name, the distribution the data was drawn from. That is the point of the dataset — it measures whether a model can read a scientific figure, scored against generator truth rather than against another annotator's guess.

Alongside the figures, the repo ships the raw responses of three vision-language models on a 150-figure subset, run twice: once on the clean figures and once on the same figures artificially aged to look like a scanned page from a mid-century journal.


At a glance

Figures 2,001 (1 panel each)
Categories 667 contour / 667 sky-gmm / 667 sky-real — exactly balanced
Questions 47,357 (21 per contour figure, 25 per sky figure)
Images 2,001 JPEG, RGB, ~2000–2700 px wide
Model responses 3 models x 150 figures x 2 conditions = 21,300 answered questions
Ground truth Generator-exact, not human-annotated
Download size ~8 GB of data; the full LFS history is considerably larger
License Apache 2.0 (see Provenance for the sky cutouts)

Repository layout

VQA_full/
  imgs/       vqa_NNNNNN.jpeg        2001 rendered figures
  qa_jsons/   vqa_NNNNNN_qa.json     2001 records: figure params + bounding
                                     boxes + underlying data + the questions

LMM_outputs_n150/                    RUN 1 -- the clean figures
  chatgpt_api/  vqa_NNNNNN_qa.pickle    150 per model: every question, with
  claude_haiku/                         the raw response the model gave
  gemini/

LMM_outputs_n150_archive_light/      RUN 2 -- the SAME 150 figures, aged
  aged_imgs/    vqa_NNNNNN.jpeg         the degraded images actually sent
  archive_manifest.json                 which effects fired, per figure
  chatgpt_api/  vqa_NNNNNN_qa.pickle    responses, as above
                vqa_NNNNNN_qa_archive.json   that figure's aging recipe
  claude_haiku/
  gemini/

example_data/                        3 illustrative figures shown twice: as
  imgs/      Picture_N00089.jpeg     rendered, and with every bounding box
  diags/     Picture_N00089.jpeg     drawn on in red, so the annotations can
                                     be checked by eye. Generator naming, NOT
                                     part of the 2001 indexed figures.

example_data/diags/ is the quickest way to see what is annotated: each red box in those images is one entry in the corresponding plot0 record — the plot area, every tick label, the axis labels, the colorbar, its ticks and its label. The three examples are illustrative only; they predate the vqa_NNNNNN numbering and are not part of the 2,001.


The figures

Three categories, 667 of each — the balance is deliberate, so accuracy in one category is never a function of how common it is.

Category plot0.type plot0.distribution What it is
contour contour random / linear / gmm A contour plot of synthetic data
sky-gmm image of the sky gmm A synthetic sky image: a Gaussian mixture, dressed in sky coordinates
sky-real image of the sky sky A real survey cutout, plotted the same way

The last two are the interesting pair: they are rendered identically and are distinguishable only by the structure of the data itself. A model that answers the provenance question correctly has to tell a real astronomical field from a plausible-looking fake.

Everything else about each figure is randomised independently — plot style, colormap, dpi, aspect ratio, font sizes, colorbar placement, tick formats, label text, coordinate epoch. No two figures share a look.

Figure ids carry no information

Ids are shuffled. vqa_000002 being a real cutout and vqa_000001 being synthetic is an accident of ordering, and nothing in the filename encodes the category. This is intentional: the generator's own filenames did encode the real-vs-synthetic answer, which would have leaked the Level 3 ground truth to anything that saw a path.


The questions

Questions are graded into three levels by what a reader has to do to answer:

Level What it asks for Examples
Level 1 Things displayed on the figure — readable directly off the page plot style, colormap, axis limits, tick values, titles, coordinate epoch
Level 2 Things requiring estimation from the rendered data min/max/median/mean of the colour axis, angular field width and height
Level 3 Things requiring inference about how the data was made which distribution generated it; real sky vs. Gaussian mixture; pixel scale

Level 3 is the hard tier and the reason the dataset exists. Level 2 is where answers are most often wrong by orders of magnitude rather than by a little.

Full question inventory

Asked of every figure (10 figure-level, Level 1):

Name Question
plot style What is the plot style used in this figure?
colormap What is the colormap that was used in this figure?
aspect ratio What is the aspect ratio of this figure?
titles What is the title of the plot in this figure?
xlabels / ylabels What is the x/y-axis title of the plot in this figure?
xtick values / ytick values What are the values for each of the tick marks on the x/y-axis?
plot types What is the plot type of the plot in this figure? (open-ended)
plot types (list) Same question, but offering [contour, image of the sky]

plot types is asked twice on purpose — once open-ended and once as a multiple choice. Only the list variant measures recognition; the open-ended one scores near zero for every model because they answer things like "astronomical survey image", which never string-matches the ground truth. Score the list variant unless you specifically want to measure phrasing.

Contour figures only (11 plot-level):

Level Name Question
1 image or lines (list) Is the contour plot drawn as an image, lines, or both?
1 minimum/maximum x axis limit Lower/upper limit of the x axis (not of the data)
1 minimum/maximum y axis limit Lower/upper limit of the y axis
2 minimum/maximum/median/mean color Statistics of the colour-axis data
3 distribution-color + list) Which distribution generated the colour data? [random, linear, gaussian mixture model]
3 distribution-x/y + list) Same, for the x/y-plane

Sky figures only (15 plot-level):

Level Name Question
1 epoch Which coordinate epoch is on the axis labels? ("none" if absent)
1 finest unit right ascension Smallest unit on the RA ticks: [hours, minutes, seconds]
1 finest unit declination Smallest unit on the Dec ticks: [degrees, arcminutes, arcseconds]
1 minimum/maximum right ascension axis limit RA axis limits
1 minimum/maximum declination axis limit Dec axis limits
2 field width / field height Angular size of the sky region, along RA / Dec
2 minimum/maximum/median/mean color Statistics of the colour-axis data
3 distribution-image + list) [gaussian mixture model, real image of the sky]the provenance question
3 pixel scale Angular size one pixel spans, in arcseconds

Totals: 21 questions per contour figure, 25 per sky figure — 667 x 21 + 1334 x 25 = 47,357.


What a qa_json contains

Each file in VQA_full/qa_jsons/ is a JSON-encoded string, so it takes two decodes:

import json
with open('VQA_full/qa_jsons/vqa_000001_qa.json') as f:
    record = json.loads(json.load(f))   # note: twice
record.keys()   # ['figure', 'plot0', 'VQA', 'vqa_id']
Key Contents
figure Figure-wide render parameters: dpi, figsize, aspect ratio, plot style, color map, font sizes, facecolor, pixel dimensions
plot0 The single panel — its type, distribution, the underlying data, and pixel bounding boxes for the plot area, every tick label, axis labels, title, colorbar and colorbar ticks
VQA The questions, as VQA[level][kind][name]
vqa_id e.g. "vqa_000001" — matches the filename and the image

Because plot0 carries bounding boxes for every element, this doubles as labelled training data for figure-element detection, not just VQA.

Reading the questions — mind the nesting

The two question kinds nest differently, which is the single most common thing to get wrong:

VQA['Level 1']['Figure-level questions']['plot style']          # -> the entry
VQA['Level 2']['Plot-level questions']['mean color']['plot0']   # -> the entry
                                                     ^^^^^^^^

Plot-level questions are keyed by panel underneath the question name. Every figure in this release has exactly one panel (plot0), so it is easy to miss — and code that skips the panel level will silently return a dict instead of a question. A safe accessor:

def question_entries(node):
    """Yield (panel, entry) whether or not the question is per-panel."""
    if 'Q' in node:                 # figure-level: the entry itself
        yield None, node
    else:                           # plot-level: {panel: entry}
        for panel, entry in node.items():
            yield panel, entry

Each entry holds the prompt in assembled and component form:

Field Meaning
Q The full assembled prompt: persona + context + question + format
A Ground truth, from the generator. Always a dict here, e.g. {"plot style": "seaborn-v0_8-muted"}
persona e.g. "You are a helpful assistant that can analyze images."
context Framing, and the multiple-choice list when one is offered
question The question alone
format The requested JSON answer shape, e.g. {"plot style":""}

The format field is worth using rather than hard-coding key names: it states exactly which JSON key the model was asked to fill, so parsing can be driven from the prompt instead of drifting from it.


The model runs

Three low-tier vision models, the same 150 figures (50 per category, vqa_000001vqa_000157), asked all 25/21 of each figure's questions:

Directory Model id as run
chatgpt_api/ gpt-5.4-nano-2026-03-17
claude_haiku/ claude-haiku-4-5
gemini/ gemini-3.5-flash-lite

Both runs use identical figures and identical questions; they differ only in whether the image was aged before being sent. Any gap between them is an effect of the degradation, not of the sample.

Reading a response pickle

import pickle
with open('LMM_outputs_n150/gemini/vqa_000001_qa.pickle', 'rb') as f:
    qa_list, model_id = pickle.load(f)

qa_list[0]['question']     # what was asked
qa_list[0]['A']            # ground truth
qa_list[0]['raw answer']   # what the model said, verbatim

Fields present for every model: Q, A, Level, type, persona, context, question, format, reasoning, prompt, Response, Response String, raw answer.

Provider-specific extras: fac (GPT); usage and system prompt (Claude, Gemini); Response (raw) (Gemini).

Unpickling needs the vendor SDKs. Each pickle stores the provider's own response object next to the text, so a plain pickle.load raises ModuleNotFoundError without anthropic / google-genai installed. If you only want the text — and the fields above are all plain strings — substitute a stub for unknown classes:

class _Stub:
    def __init__(self, *a, **k): pass
    def __setstate__(self, s): self.__dict__.update(s if isinstance(s, dict) else {})

class TolerantUnpickler(pickle.Unpickler):
    def find_class(self, module, name):
        try:
            return super().find_class(module, name)
        except (ImportError, AttributeError):
            return type(name, (_Stub,), {})

The archive-light condition

LMM_outputs_n150_archive_light/ asks the same questions of the same figures after degrading each one to resemble a scanned page from an old journal — ink bleed and mottling, paper colouring and texture, stains, folding, subtle noise, JPEG artefacts, and a 2-in-3 chance of being converted to grayscale.

The aging is reproducible and shared across models: each figure's random draw is seeded from its own id, so GPT, Claude and Gemini all saw the byte-identical aged image. Comparing the three models would otherwise be confounded by each having been shown a differently-damaged page.

archive_manifest.json records the draw for all 150 figures; the same record is repeated per figure as <id>_qa_archive.json:

Field Meaning
preset "archive-light"
base_seed / seed Run-wide seed, and this figure's derived seed
effects Which effects fired, in order
effects_by_phase The same, split into ink / paper / post
grayscale, gray_prob Whether this figure went gray, and with what probability
page background RGB of the simulated paper

The aged images themselves are in aged_imgs/ — you do not need to re-run the aging to reproduce the experiment.


Known issues and caveats

Read these before reporting numbers.

  • Score the list variant of plot types. The open-ended variant scores near zero for all models for phrasing reasons, not vision reasons. See above.
  • Two different "distribution" questions exist. For contour figures, distribution-color and distribution-x/y over [random, linear, gaussian mixture model]. For sky figures, distribution-image over [gaussian mixture model, real image of the sky]. They are not comparable and should not be pooled — match on the full question text, not on the substring "underlying distribution".
  • Ground truth is shaped differently in the two file types. In the released qa jsons A is always a dict (all 47,357 of them). In the run pickles it is a dict for figure-level questions but unwrapped to a bare str or float for plot-level ones. Each question is consistent in its own type, but code that reads both file types must handle both shapes.
  • Models rename the answer key. Claude commonly answers plot_types when asked for plot types. Match keys up to spacing, underscores and case, or you will score a correct answer as a parse failure.
  • A small number of responses are unparsable — about 0.14% (30 of 21,300), split between prose-instead-of-JSON and valid JSON containing an explicit null refusal. Decide deliberately whether those count as wrong or as excluded; the two failure modes arguably differ.
  • Level 2 answers are often wrong by orders of magnitude. Bounded metrics (e.g. sSMAPE) saturate here. log10(predicted/true) is the more informative summary for the angular-size and colour-statistic questions.
  • The dataset viewer is disabled (viewer: false). The JSON records are deeply nested rather than tabular, so the viewer cannot render them meaningfully.

Provenance and attribution

The sky-real figures are plotted from real survey cutouts retrieved via NASA SkyView. Across the 667 real-sky figures: 659 distinct astronomical objects, drawn from 14 SkyView survey categories — predominantly optical (OtherOptical, Optical:DSS), with infrared (IRAS, WISE, 2MASS, AKARI, Planck, WMAP&COBE, UKIDSS), radio (Radio:MHz, Radio:GHz, GLEAM) and X-ray (ROSATDiffuse) representative among them.

Each figure's plot0.data['data params']['sky image params'] keeps the full provenance: source filename, the original FITS header, survey, telescope, instrument, object, bunit, and the ADS-style bibcode of the paper the field was taken from (411 distinct source papers).

If you publish work using the sky-real figures, you inherit the acknowledgement obligations of the underlying surveys, which are not waived by this repo's Apache 2.0 license. In particular SkyView asks to be cited, and the Digitized Sky Surveys — the largest single source here — carry their own acknowledgement requirements from STScI and the originating plate collections. Check the terms for the specific surveys in your subset; the metadata above tells you which those are.

The Apache 2.0 license applies to the dataset as assembled: the synthetic figures, the generated questions and ground truth, and the model responses.

Generating code

The figures and questions are produced by the SkyImagesWASP2026 repository.

Citation

tbd
Downloads last month
210