|
Download README.md from CodeHima/TOS_DatasetV3: direct link, hf CLI and curl.
- Browser
- Download file 3.07 kB
-
https://huggingface.co/datasets/CodeHima/TOS_DatasetV3/resolve/main/README.md
- Command line
-
hf download hf://datasets/CodeHima/TOS_DatasetV3/README.md
-
curl -L -o README.md https://huggingface.co/datasets/CodeHima/TOS_DatasetV3/resolve/main/README.md
3.07 kB
| language: | |
| - en | |
| dataset_info: | |
| features: | |
| - name: sentence | |
| dtype: string | |
| - name: unfairness_level | |
| dtype: string | |
| splits: | |
| - name: train | |
| num_bytes: 1221140 | |
| num_examples: 7945 | |
| - name: validation | |
| num_bytes: 151516 | |
| num_examples: 1050 | |
| - name: test | |
| num_bytes: 155836 | |
| num_examples: 1045 | |
| download_size: 708831 | |
| dataset_size: 1528492 | |
| configs: | |
| - config_name: default | |
| data_files: | |
| - split: train | |
| path: data/train-* | |
| - split: validation | |
| path: data/validation-* | |
| - split: test | |
| path: data/test-* | |
| # TOS_DatasetV3 | |
| ## Dataset Description | |
| TOS_DatasetV3 is a dataset designed for analyzing the unfairness of terms of service (ToS) clauses. It includes sentences from various terms of service agreements categorized into three unfairness levels: **clearly_fair**, **potentially_unfair**, and **clearly_unfair**. This dataset aims to aid in the development of models that can assess the fairness of legal documents. | |
| ## Dataset Structure | |
| The dataset consists of the following columns: | |
| - `sentence`: A string representing a sentence from a terms of service document. | |
| - `unfairness_level`: A string indicating the unfairness classification of the sentence. The possible values are: | |
| - `clearly_fair` | |
| - `potentially_unfair` | |
| - `clearly_unfair` | |
| ### Features | |
| - **Features:** | |
| - name: sentence | |
| dtype: string | |
| - name: unfairness_level | |
| dtype: string | |
| ## Splits | |
| The dataset is divided into three splits: | |
| - **Train**: Used for training models (7946 examples). | |
| - **Validation**: Used for validating model performance during training (1050 examples). | |
| - **Test**: Used for testing model performance after training (1045 examples). | |
| ## Class Split | |
|  | |
| ### Example | |
| Here's a sample from the dataset: | |
| | sentence | unfairness_level | | |
| |-----------------------------------|---------------------| | |
| | "You agree to the terms." | clearly_fair | | |
| | "We reserve the right to change." | potentially_unfair | | |
| | "No refunds will be issued." | clearly_unfair | | |
| ## Usage | |
| This dataset can be used for various natural language processing tasks, particularly in the context of legal text analysis, fairness assessment, and model training for detecting unfair terms in contracts. | |
| ## How to Load the Dataset | |
| You can load the dataset using the `datasets` library from Hugging Face: | |
| ```python | |
| from datasets import load_dataset | |
| dataset = load_dataset("TOS_DatasetV3") | |
| ``` | |
| ## Dataset Generation | |
|  | |
| ## License | |
| This dataset is licensed under the MIT License. Please see the LICENSE file for more details. | |
| ## Citation | |
| If you use this dataset in your research, please cite it as follows: | |
| ``` | |
| @dataset{TOS_DatasetV3, | |
| author = {Himanshu Mohanty}, | |
| title = {TOS_DatasetV3}, | |
| year = {2024}, | |
| url = {https://huggingface.co/datasets/TOS_DatasetV3} | |
| } |