Datasets:
The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
Error code: DatasetGenerationError
Exception: TypeError
Message: Couldn't cast array of type
struct<corpus_id: string, query_id: string, raw: string, score: double, annotation: null>
to
{'annotation': Json(decode=True), 'explanation': Value('string'), 'filename': Value('string'), 'has_ground_truth': Value('bool'), 'program': Value('string'), 'program_re': Value('string'), 'raw': Json(decode=True), 'turn_name': Value('string')}
Traceback: Traceback (most recent call last):
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1816, in _prepare_split_single
for key, table in generator:
^^^^^^^^^
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 613, in wrapped
for item in generator(*args, **kwargs):
~~~~~~~~~^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 343, in _generate_tables
self._cast_table(pa_table, json_field_paths=json_field_paths),
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 132, in _cast_table
pa_table = table_cast(pa_table, self.info.features.arrow_schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2369, in table_cast
return cast_table_to_schema(table, schema)
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2303, in cast_table_to_schema
cast_array_to_feature(
~~~~~~~~~~~~~~~~~~~~~^
table[name] if name in table_column_names else pa.array([None] * len(table), type=schema.field(name).type),
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
feature,
^^^^^^^^
)
^
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 1852, in wrapper
return pa.chunked_array([func(chunk, *args, **kwargs) for chunk in array.chunks])
~~~~^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/table.py", line 2149, in cast_array_to_feature
raise TypeError(f"Couldn't cast array of type\n{_short_str(array.type)}\nto\n{_short_str(feature)}")
TypeError: Couldn't cast array of type
struct<corpus_id: string, query_id: string, raw: string, score: double, annotation: null>
to
{'annotation': Json(decode=True), 'explanation': Value('string'), 'filename': Value('string'), 'has_ground_truth': Value('bool'), 'program': Value('string'), 'program_re': Value('string'), 'raw': Json(decode=True), 'turn_name': Value('string')}
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
~~~~~~~~~~~~~~~~~~~~~~~~~^
builder, max_dataset_size_bytes=max_dataset_size_bytes
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1683, in _prepare_split
for job_id, done, content in self._prepare_split_single(
~~~~~~~~~~~~~~~~~~~~~~~~~~^
gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
):
^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1869, in _prepare_split_single
raise DatasetGenerationError("An error occurred while generating the dataset") from e
datasets.exceptions.DatasetGenerationError: An error occurred while generating the datasetNeed help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
dataset string | id string | input string | instruction string | metadata dict | output string | schema_version string | split string |
|---|---|---|---|---|---|---|---|
ConvFinQA | Single_JKHY/2009/page_28.pdf-3-qa | Pre text: 26 | 2009 annual report in fiscal 2008 , revenues in the credit union systems and services business segment increased 14% ( 14 % ) from fiscal 2007 . all revenue components within the segment experienced growth during fiscal 2008 . license revenue generated the largest dollar growth in revenue as episys ae , ... | what was the percentage change in the net cash from operating activities from 2008 to 2009 | {
"annotation": {
"amt_post_text": "year ended june 30 , cash provided by operations increased $ 25587 to $ 206588 for the fiscal year ended june 30 , 2009 as compared to $ 181001 for the fiscal year ended june 30 , 2008 . this increase is primarily attributable to a decrease in receivables compared to the same p... | 14.1% | 1.0 | train |
ConvFinQA | Single_RSG/2008/page_114.pdf-2-qa | Pre text: substantially all of the goodwill and other intangible assets recorded related to the acquisition of allied are not deductible for tax purposes . pro forma information the consolidated financial statements presented for republic include the operating results of allied from the date of the acquisition . the fo... | what was the percent of the growth in the revenues from 2007 to 2008 | {
"annotation": {
"amt_post_text": "the above unaudited pro forma financial information includes adjustments for amortization of identifiable intangible assets , accretion of discounts to fair value associated with debt , environmental , self-insurance and other liabilities , accretion of capping , closure and po... | 1.3% | 1.0 | train |
ConvFinQA | Single_AAPL/2002/page_23.pdf-1-qa | Pre text: in a new business model such as the retail segment is inherently risky , particularly in light of the significant investment involved , the current economic climate , and the fixed nature of a substantial portion of the retail segment's operating expenses . results for this segment are dependent upon a number... | what was the percentage change in net sales from 2000 to 2001? | {
"annotation": {
"amt_post_text": ".",
"amt_pre_text": "in a new business model such as the retail segment is inherently risky , particularly in light of the significant investment involved , the current economic climate , and the fixed nature of a substantial portion of the retail segment's operating expens... | -32% | 1.0 | train |
ConvFinQA | Single_UPS/2009/page_33.pdf-2-qa | Pre text: ( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph and related information shall not be deemed 201csoliciting... | what was the difference in percentage cumulative return on investment for united parcel service inc . compared to the s&p 500 index for the five year period ended 12/31/09? | {
"annotation": {
"amt_post_text": ".",
"amt_pre_text": "( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph a... | -26.16% | 1.0 | train |
ConvFinQA | Double_UPS/2009/page_33.pdf-qa_0 | Pre text: ( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph and related information shall not be deemed 201csoliciting... | what is the roi of an investment in ups in 2004 and sold in 2006? | {
"annotation": {
"amt_post_text": ".",
"amt_pre_text": "( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph a... | -8.9% | 1.0 | train |
ConvFinQA | Double_UPS/2009/page_33.pdf-qa_1 | Pre text: ( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph and related information shall not be deemed 201csoliciting... | what was the difference in percentage cumulative return on investment for united parcel service inc . compared to the s&p 500 index for the five year period ended 12/31/09? | {
"annotation": {
"amt_post_text": ".",
"amt_pre_text": "( 1 ) includes shares repurchased through our publicly announced share repurchase program and shares tendered to pay the exercise price and tax withholding on employee stock options . shareowner return performance graph the following performance graph a... | -26.16% | 1.0 | train |
ConvFinQA | Single_CE/2010/page_134.pdf-2-qa | Pre text: tax returns for 2001 and beyond are open for examination under statute . currently , unrecognized tax benefits are not expected to change significantly over the next 12 months . 19 . stock-based and other management compensation plans in april 2009 , the company approved a global incentive plan which replaces... | what portion of the total shares subject to outstanding awards is under the 2009 global incentive plan? | {
"annotation": {
"amt_post_text": "upon the termination of a participant 2019s employment with the company by reason of death or disability or by the company without cause ( as defined in the respective award agreements ) , an award in amount equal to ( i ) the value of the award granted multiplied by ( ii ) a f... | 70.1% | 1.0 | train |
ConvFinQA | Single_JPM/2013/page_104.pdf-2-qa | Pre text: management 2019s discussion and analysis 110 jpmorgan chase & co./2013 annual report 2012 compared with 2011 net loss was $ 2.0 billion , compared with a net income of $ 919 million in the prior year . private equity reported net income of $ 292 million , compared with net income of $ 391 million in the prior... | what was the percentage increase in litigation reserves in 2012? | {
"annotation": {
"amt_post_text": "( a ) period-end investment securities included held-to-maturity balance of $ 24.0 billion at december 31 , 2013 . held-to-maturity balances for the other periods were not material. .",
"amt_pre_text": "management 2019s discussion and analysis 110 jpmorgan chase & co./2013 ... | 15.6% | 1.0 | train |
ConvFinQA | Double_MAS/2012/page_92.pdf-qa_0 | Pre text: masco corporation notes to consolidated financial statements ( continued ) t . other commitments and contingencies litigation . we are subject to claims , charges , litigation and other proceedings in the ordinary course of our business , including those arising from or related to contractual matters , intell... | what was the percent of the change in the company 2019s warranty liability from 2011 to 2012 | {
"annotation": {
"amt_post_text": "investments . with respect to the company 2019s investments in private equity funds , the company had , at december 31 , 2012 , commitments to contribute up to $ 19 million of additional capital to such funds representing the company 2019s aggregate capital commitment to such f... | 15.7% | 1.0 | train |
ConvFinQA | Double_MAS/2012/page_92.pdf-qa_1 | "Pre text: masco corporation notes to consolidated financial statements ( continued ) t . other comm(...TRUNCATED) | what was the percentage change in the company's warranty liability from 2011 to 2012? | {"annotation":{"amt_post_text":"investments . with respect to the company 2019s investments in priva(...TRUNCATED) | 16% | 1.0 | train |
FinGemma Unified Financial Dataset
The FinGemma Dataset is a production-quality, aggregated, and fully standardized instruction fine-tuning dataset designed primarily for training and evaluating financial large language models (such as FT_Gemma) with advanced document-grounded question answering (QA) and financial reasoning capabilities. It consolidates 8 premium financial datasets spanning arithmetic calculations, tabular comprehension, red-teaming safety, and sentiment analysis.
1. Overview & Dataset Description
This unified dataset provides 63,820 standardized question-answering, multi-turn dialogue, safety, and sentiment examples. Every row is normalized into a standard, deterministic layout containing:
- Instruction: Direct question, prompt, or safety query.
- Input: Accompanying structured context, including preceding texts, aligned markdown-rendered tables, and document metadata.
- Output: Validated ground truth answer (if available; unlabeled private test sets are gracefully mapped with blank strings).
- Metadata: Highly detailed trace metrics (formulas, original programs, sectors, filenames) as well as the complete, preserved original raw record to prevent any data loss during preprocessing.
2. Dataset Statistics
Consolidated Master Splits
These represent the merged, ready-to-train datasets consolidated from all 8 raw sources:
| Split | Total Samples | Format |
|---|---|---|
| test | 7,979 | JSONL |
| train | 49,907 | JSONL |
| validation | 5,934 | JSONL |
Aggregated Source Subsets
| Source Dataset | Domain / Task | Version | Original Source Repo / URL |
|---|---|---|---|
| FinQA | financial-reasoning | emnlp-2021 |
[GitHub] czyssrs/FinQA |
| ConvFinQA | conversational-finance-reasoning | emnlp-2022 |
[Direct Link] https://raw.githubusercontent.com/czyssrs/Con... |
| TAT-QA | table-text-qa | acl-2021 |
[GitHub] NExTplusplus/TAT-QA |
| FinanceBench | open-book-financial-qa | 2023-open-source |
[HF] PatronusAI/financebench |
| FinancialPhraseBank | financial-sentiment | sentences_allagree |
[Direct Link] https://huggingface.co/datasets/takala/financ... |
| FiQA | financial-opinion-qa | www-2018 |
[HF] mteb/fiqa |
| DocFinQA | document-financial-qa | hf |
[HF] kensho/DocFinQA |
| FinRED | financial-safety | hf |
[HF] datumo/FinRED |
3. Source Datasets & Taxonomy
- FinQA: FinQA numerical reasoning dataset
- ConvFinQA: ConvFinQA conversational finance reasoning dataset
- TAT-QA: TAT-QA hybrid tabular and textual QA dataset
- FinanceBench: FinanceBench open-source financial QA sample
- FinancialPhraseBank: Financial PhraseBank sentiment dataset
- FiQA: FiQA 2018 opinion QA challenge dataset
- DocFinQA: DocFinQA financial document QA dataset
- FinRED: FinRED financial red-teaming benchmark
4. Unified Schema (v1.0)
Each record in the .jsonl split files strictly adheres to the following unified, flat JSON schema:
{
"schema_version": "1.0",
"dataset": "FinQA",
"split": "train",
"id": "ADI/2009/page_49.pdf-1",
"instruction": "what is the interest expense in 2009?",
"input": "Preceding text:\n...\n\nTable:\n...\n\nFollowing text:\n...",
"output": "380",
"metadata": {
"filename": "ADI/2009/page_49.pdf",
"program": "divide(100, 100), divide(3.8, #0)",
"raw": { ... }
}
}
5. Repository Directory Structure
The repository contains both consolidated master split files and independent subfolders for granular Loading/Filtering:
datasets/processed/
βββ train.jsonl # Consolidated master training splits
βββ validation.jsonl # Consolidated master validation splits
βββ test.jsonl # Consolidated master test splits
βββ FinQA/ # FinQA subset splits
βββ ConvFinQA/ # ConvFinQA subset splits
βββ TAT-QA/ # TAT-QA subset splits
βββ FinanceBench/ # FinanceBench subset splits
βββ FinancialPhraseBank/ # FinancialPhraseBank subset splits
βββ FiQA/ # FiQA subset splits
βββ DocFinQA/ # DocFinQA subset splits
βββ FinRED/ # FinRED subset splits
6. Intended Use
This dataset is designed specifically for:
- Fine-tuning Large Language Models on complex, multi-modal financial math and text.
- Evaluating mathematical, conversational, safety, and document understanding capacities.
- Aligning models to perform instruction-following, document comprehension, and complex financial calculations.
7. Limitations & Edge Cases
- Unlabeled Private Tests: Official private test sets for
FinQA,ConvFinQA, andTAT-QAintentionally lack ground-truth answers. These are preserved with"has_ground_truth": falseto support inference-level benchmarking. - Large Context Sizes:
DocFinQAcontains long text Context records (often exceeding 330 KB). Attention sliding windows or robust retrieval-based chunking is recommended during training.
8. License & Terms
This merged collection is published under the Apache 2.0 License where wrapping schemas and code are concerned. Please refer to the licenses of the original datasets (e.g. Creative Commons, non-commercial licenses) for their respective data inputs.
9. Generation Info
- Dataset Schema Version:
1.0 - Last Generated:
2026-07-09 08:10:59 UTC
- Downloads last month
- 82