Dataset Preview
Duplicate
The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
The dataset generation failed
Error code:   DatasetGenerationError
Exception:    TypeError
Message:      Couldn't cast array of type
struct<corpus_id: string, query_id: string, raw: string, score: double, annotation: null>
to
{'annotation': Json(decode=True), 'explanation': Value('string'), 'filename': Value('string'), 'has_ground_truth': Value('bool'), 'program': Value('string'), 'program_re': Value('string'), 'raw': Json(decode=True), 'turn_name': Value('string')}
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1816, in _prepare_split_single
                  for key, table in generator:
                                    ^^^^^^^^^
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 613, in wrapped
                  for item in generator(*args, **kwargs):
                              ~~~~~~~~~^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 343, in _generate_tables
                  self._cast_table(pa_table, json_field_paths=json_field_paths),
                  ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 132, in _cast_table
                  pa_table = table_cast(pa_table, self.info.features.arrow_schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2369, in table_cast
                  return cast_table_to_schema(table, schema)
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2303, in cast_table_to_schema
                  cast_array_to_feature(
                  ~~~~~~~~~~~~~~~~~~~~~^
                      table[name] if name in table_column_names else pa.array([None] * len(table), type=schema.field(name).type),
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                      feature,
                      ^^^^^^^^
                  )
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 1852, in wrapper
                  return pa.chunked_array([func(chunk, *args, **kwargs) for chunk in array.chunks])
                                           ~~~~^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2149, in cast_array_to_feature
                  raise TypeError(f"Couldn't cast array of type\n{_short_str(array.type)}\nto\n{_short_str(feature)}")
              TypeError: Couldn't cast array of type
              struct<corpus_id: string, query_id: string, raw: string, score: double, annotation: null>
              to
              {'annotation': Json(decode=True), 'explanation': Value('string'), 'filename': Value('string'), 'has_ground_truth': Value('bool'), 'program': Value('string'), 'program_re': Value('string'), 'raw': Json(decode=True), 'turn_name': Value('string')}
              
              The above exception was the direct cause of the following exception:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
                  parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
                                                                        ~~~~~~~~~~~~~~~~~~~~~~~~~^
                      builder, max_dataset_size_bytes=max_dataset_size_bytes
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  )
                  ^
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
                  builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
                  ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1683, in _prepare_split
                  for job_id, done, content in self._prepare_split_single(
                                               ~~~~~~~~~~~~~~~~~~~~~~~~~~^
                      gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  ):
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1869, in _prepare_split_single
                  raise DatasetGenerationError("An error occurred while generating the dataset") from e
              datasets.exceptions.DatasetGenerationError: An error occurred while generating the dataset

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

dataset
string
id
string
input
string
instruction
string
metadata
dict
output
string
schema_version
string
split
string
ConvFinQA
Single_JKHY/2009/page_28.pdf-3-qa
Pre text: 26 | 2009 annual report in fiscal 2008 , revenues in the credit union systems and services business segment increased 14% ( 14 % ) from fiscal 2007 . all revenue components within the segment experienced growth during fiscal 2008 . license revenue generated the largest dollar growth in revenue as episys ae , ...
what was the percentage change in the net cash from operating activities from 2008 to 2009
{ "annotation": { "amt_post_text": "year ended june 30 , cash provided by operations increased $ 25587 to $ 206588 for the fiscal year ended june 30 , 2009 as compared to $ 181001 for the fiscal year ended june 30 , 2008 . this increase is primarily attributable to a decrease in receivables compared to the same p...
14.1%
1.0
train
ConvFinQA
Single_RSG/2008/page_114.pdf-2-qa
Pre text: substantially all of the goodwill and other intangible assets recorded related to the acquisition of allied are not deductible for tax purposes . pro forma information the consolidated financial statements presented for republic include the operating results of allied from the date of the acquisition . the fo...
what was the percent of the growth in the revenues from 2007 to 2008
{ "annotation": { "amt_post_text": "the above unaudited pro forma financial information includes adjustments for amortization of identifiable intangible assets , accretion of discounts to fair value associated with debt , environmental , self-insurance and other liabilities , accretion of capping , closure and po...
1.3%
1.0
train
ConvFinQA
Single_AAPL/2002/page_23.pdf-1-qa
Pre text: in a new business model such as the retail segment is inherently risky , particularly in light of the significant investment involved , the current economic climate , and the fixed nature of a substantial portion of the retail segment's operating expenses . results for this segment are dependent upon a number...
what was the percentage change in net sales from 2000 to 2001?
{ "annotation": { "amt_post_text": ".", "amt_pre_text": "in a new business model such as the retail segment is inherently risky , particularly in light of the significant investment involved , the current economic climate , and the fixed nature of a substantial portion of the retail segment's operating expens...
-32%
1.0
train
ConvFinQA
Single_UPS/2009/page_33.pdf-2-qa
Pre text: ( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph and related information shall not be deemed 201csoliciting...
what was the difference in percentage cumulative return on investment for united parcel service inc . compared to the s&p 500 index for the five year period ended 12/31/09?
{ "annotation": { "amt_post_text": ".", "amt_pre_text": "( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph a...
-26.16%
1.0
train
ConvFinQA
Double_UPS/2009/page_33.pdf-qa_0
Pre text: ( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph and related information shall not be deemed 201csoliciting...
what is the roi of an investment in ups in 2004 and sold in 2006?
{ "annotation": { "amt_post_text": ".", "amt_pre_text": "( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph a...
-8.9%
1.0
train
ConvFinQA
Double_UPS/2009/page_33.pdf-qa_1
Pre text: ( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph and related information shall not be deemed 201csoliciting...
what was the difference in percentage cumulative return on investment for united parcel service inc . compared to the s&p 500 index for the five year period ended 12/31/09?
{ "annotation": { "amt_post_text": ".", "amt_pre_text": "( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph a...
-26.16%
1.0
train
ConvFinQA
Single_CE/2010/page_134.pdf-2-qa
Pre text: tax returns for 2001 and beyond are open for examination under statute . currently , unrecognized tax benefits are not expected to change significantly over the next 12 months . 19 . stock-based and other management compensation plans in april 2009 , the company approved a global incentive plan which replaces...
what portion of the total shares subject to outstanding awards is under the 2009 global incentive plan?
{ "annotation": { "amt_post_text": "upon the termination of a participant 2019s employment with the company by reason of death or disability or by the company without cause ( as defined in the respective award agreements ) , an award in amount equal to ( i ) the value of the award granted multiplied by ( ii ) a f...
70.1%
1.0
train
ConvFinQA
Single_JPM/2013/page_104.pdf-2-qa
Pre text: management 2019s discussion and analysis 110 jpmorgan chase & co./2013 annual report 2012 compared with 2011 net loss was $ 2.0 billion , compared with a net income of $ 919 million in the prior year . private equity reported net income of $ 292 million , compared with net income of $ 391 million in the prior...
what was the percentage increase in litigation reserves in 2012?
{ "annotation": { "amt_post_text": "( a ) period-end investment securities included held-to-maturity balance of $ 24.0 billion at december 31 , 2013 . held-to-maturity balances for the other periods were not material. .", "amt_pre_text": "management 2019s discussion and analysis 110 jpmorgan chase & co./2013 ...
15.6%
1.0
train
ConvFinQA
Double_MAS/2012/page_92.pdf-qa_0
Pre text: masco corporation notes to consolidated financial statements ( continued ) t . other commitments and contingencies litigation . we are subject to claims , charges , litigation and other proceedings in the ordinary course of our business , including those arising from or related to contractual matters , intell...
what was the percent of the change in the company 2019s warranty liability from 2011 to 2012
{ "annotation": { "amt_post_text": "investments . with respect to the company 2019s investments in private equity funds , the company had , at december 31 , 2012 , commitments to contribute up to $ 19 million of additional capital to such funds representing the company 2019s aggregate capital commitment to such f...
15.7%
1.0
train
ConvFinQA
Double_MAS/2012/page_92.pdf-qa_1
"Pre text: masco corporation notes to consolidated financial statements ( continued ) t . other comm(...TRUNCATED)
what was the percentage change in the company's warranty liability from 2011 to 2012?
{"annotation":{"amt_post_text":"investments . with respect to the company 2019s investments in priva(...TRUNCATED)
16%
1.0
train
End of preview.

FinGemma Unified Financial Dataset

The FinGemma Dataset is a production-quality, aggregated, and fully standardized instruction fine-tuning dataset designed primarily for training and evaluating financial large language models (such as FT_Gemma) with advanced document-grounded question answering (QA) and financial reasoning capabilities. It consolidates 8 premium financial datasets spanning arithmetic calculations, tabular comprehension, red-teaming safety, and sentiment analysis.


1. Overview & Dataset Description

This unified dataset provides 63,820 standardized question-answering, multi-turn dialogue, safety, and sentiment examples. Every row is normalized into a standard, deterministic layout containing:

  • Instruction: Direct question, prompt, or safety query.
  • Input: Accompanying structured context, including preceding texts, aligned markdown-rendered tables, and document metadata.
  • Output: Validated ground truth answer (if available; unlabeled private test sets are gracefully mapped with blank strings).
  • Metadata: Highly detailed trace metrics (formulas, original programs, sectors, filenames) as well as the complete, preserved original raw record to prevent any data loss during preprocessing.

2. Dataset Statistics

Consolidated Master Splits

These represent the merged, ready-to-train datasets consolidated from all 8 raw sources:

Split Total Samples Format
test 7,979 JSONL
train 49,907 JSONL
validation 5,934 JSONL

Aggregated Source Subsets

Source Dataset Domain / Task Version Original Source Repo / URL
FinQA financial-reasoning emnlp-2021 [GitHub] czyssrs/FinQA
ConvFinQA conversational-finance-reasoning emnlp-2022 [Direct Link] https://raw.githubusercontent.com/czyssrs/Con...
TAT-QA table-text-qa acl-2021 [GitHub] NExTplusplus/TAT-QA
FinanceBench open-book-financial-qa 2023-open-source [HF] PatronusAI/financebench
FinancialPhraseBank financial-sentiment sentences_allagree [Direct Link] https://huggingface.co/datasets/takala/financ...
FiQA financial-opinion-qa www-2018 [HF] mteb/fiqa
DocFinQA document-financial-qa hf [HF] kensho/DocFinQA
FinRED financial-safety hf [HF] datumo/FinRED

3. Source Datasets & Taxonomy

  • FinQA: FinQA numerical reasoning dataset
  • ConvFinQA: ConvFinQA conversational finance reasoning dataset
  • TAT-QA: TAT-QA hybrid tabular and textual QA dataset
  • FinanceBench: FinanceBench open-source financial QA sample
  • FinancialPhraseBank: Financial PhraseBank sentiment dataset
  • FiQA: FiQA 2018 opinion QA challenge dataset
  • DocFinQA: DocFinQA financial document QA dataset
  • FinRED: FinRED financial red-teaming benchmark

4. Unified Schema (v1.0)

Each record in the .jsonl split files strictly adheres to the following unified, flat JSON schema:

{
  "schema_version": "1.0",
  "dataset": "FinQA",
  "split": "train",
  "id": "ADI/2009/page_49.pdf-1",
  "instruction": "what is the interest expense in 2009?",
  "input": "Preceding text:\n...\n\nTable:\n...\n\nFollowing text:\n...",
  "output": "380",
  "metadata": {
    "filename": "ADI/2009/page_49.pdf",
    "program": "divide(100, 100), divide(3.8, #0)",
    "raw": { ... }
  }
}

5. Repository Directory Structure

The repository contains both consolidated master split files and independent subfolders for granular Loading/Filtering:

datasets/processed/
β”œβ”€β”€ train.jsonl             # Consolidated master training splits
β”œβ”€β”€ validation.jsonl        # Consolidated master validation splits
β”œβ”€β”€ test.jsonl              # Consolidated master test splits
β”œβ”€β”€ FinQA/                  # FinQA subset splits
β”œβ”€β”€ ConvFinQA/              # ConvFinQA subset splits
β”œβ”€β”€ TAT-QA/                 # TAT-QA subset splits
β”œβ”€β”€ FinanceBench/           # FinanceBench subset splits
β”œβ”€β”€ FinancialPhraseBank/    # FinancialPhraseBank subset splits
β”œβ”€β”€ FiQA/                   # FiQA subset splits
β”œβ”€β”€ DocFinQA/               # DocFinQA subset splits
└── FinRED/                 # FinRED subset splits

6. Intended Use

This dataset is designed specifically for:

  • Fine-tuning Large Language Models on complex, multi-modal financial math and text.
  • Evaluating mathematical, conversational, safety, and document understanding capacities.
  • Aligning models to perform instruction-following, document comprehension, and complex financial calculations.

7. Limitations & Edge Cases

  • Unlabeled Private Tests: Official private test sets for FinQA, ConvFinQA, and TAT-QA intentionally lack ground-truth answers. These are preserved with "has_ground_truth": false to support inference-level benchmarking.
  • Large Context Sizes: DocFinQA contains long text Context records (often exceeding 330 KB). Attention sliding windows or robust retrieval-based chunking is recommended during training.

8. License & Terms

This merged collection is published under the Apache 2.0 License where wrapping schemas and code are concerned. Please refer to the licenses of the original datasets (e.g. Creative Commons, non-commercial licenses) for their respective data inputs.


9. Generation Info

  • Dataset Schema Version: 1.0
  • Last Generated: 2026-07-09 08:10:59 UTC
Downloads last month
82