|
Download README.md from infinityofspace/python_codestyles-single-500: direct link, hf CLI and curl.
- Browser
- Download file 2.74 kB
-
https://huggingface.co/datasets/infinityofspace/python_codestyles-single-500/resolve/main/README.md
- Command line
-
hf download hf://datasets/infinityofspace/python_codestyles-single-500/README.md
-
curl -L -o README.md https://huggingface.co/datasets/infinityofspace/python_codestyles-single-500/resolve/main/README.md
2.74 kB
| configs: | |
| - config_name: default | |
| data_files: | |
| - split: train | |
| path: data/train-* | |
| - split: test | |
| path: data/test-* | |
| dataset_info: | |
| features: | |
| - name: code | |
| dtype: string | |
| - name: code_codestyle | |
| dtype: int64 | |
| - name: style_context | |
| dtype: string | |
| - name: style_context_codestyle | |
| dtype: int64 | |
| - name: label | |
| dtype: int64 | |
| splits: | |
| - name: train | |
| num_bytes: 1784386100 | |
| num_examples: 153991 | |
| - name: test | |
| num_bytes: 323920285 | |
| num_examples: 28193 | |
| download_size: 320183832 | |
| dataset_size: 2108306385 | |
| license: mit | |
| tags: | |
| - python | |
| - code-style | |
| - single | |
| size_categories: | |
| - 100K<n<1M | |
| # Dataset Card for "python_codestyles-single-500" | |
| This dataset contains negative and positive examples with python code of compliance with a code style. A positive | |
| example represents compliance with the code style (label is 1). Each example is composed of two components, the first | |
| component consists of a code that either conforms to the code style or violates it and the second component | |
| corresponding to an example code that already conforms to a code style. In total, the dataset contains `500` completely | |
| different code styles. The code styles differ in exactly one codestyle rule, which is called a `single` codestyle | |
| dataset variant. The dataset consists of a training and test group, with none of the code styles overlapping between | |
| groups. In addition, both groups contain completely different underlying codes. | |
| The examples contain source code from the following repositories: | |
| | repository | tag or commit | | |
| |:-----------------------------------------------------------------------:|:----------------------------------------:| | |
| | [TheAlgorithms/Python](https://github.com/TheAlgorithms/Python) | f614ed72170011d2d439f7901e1c8daa7deac8c4 | | |
| | [huggingface/transformers](https://github.com/huggingface/transformers) | v4.31.0 | | |
| | [huggingface/datasets](https://github.com/huggingface/datasets) | 2.13.1 | | |
| | [huggingface/diffusers](https://github.com/huggingface/diffusers) | v0.18.2 | | |
| | [huggingface/accelerate](https://github.com/huggingface/accelerate) | v0.21.0 | | |
| You can find the corresponding code styles of the examples in the file [additional_data.json](additional_data.json). | |
| The code styles in the file are split by training and test group and the index corresponds to the class for the | |
| columns `code_codestyle` and `style_context_codestyle` in the dataset. | |
| There are 182.184 samples in total and 91.084 positive and 91.100 negative samples. |