Home

Awesome

The Pile Replication Code

The official website for the the Pile is here.

The Pile is a large, diverse, open source language modelling data set that consists of many smaller datasets combined together. The objective is to obtain text from as many modalities as possible to ensure that models trained using The Pile will have much broader generalization abilities.

This repository is for replicating or making variants of the Pile. IF YOU ARE HERE TO USE THE PILE DATASET, THIS REPO IS PROBABLY NOT WHAT YOU ARE LOOKING FOR. A copy of the Pile can be downloaded here.

ComponentRaw SizeWeightEpochsEffective SizeMean Document Size
Pile-CC227.12 GiB18.11%1.0227.12 GiB4.33 KiB
PubMed Central90.27 GiB14.40%2.0180.55 GiB30.55 KiB
Books3100.96 GiB12.07%1.5151.44 GiB538.36 KiB
OpenWebText262.77 GiB10.01%2.0125.54 GiB3.85 KiB
ArXiv56.21 GiB8.96%2.0112.42 GiB46.61 KiB
Github95.16 GiB7.59%1.095.16 GiB5.25 KiB
FreeLaw51.15 GiB6.12%1.576.73 GiB15.06 KiB
StackExchange32.20 GiB5.13%2.064.39 GiB2.16 KiB
USPTO Backgrounds22.90 GiB3.65%2.045.81 GiB4.08 KiB
PubMed Abstracts19.26 GiB3.07%2.038.53 GiB1.30 KiB
Gutenberg (PG-19)10.88 GiB2.17%2.527.19 GiB398.73 KiB
OpenSubtitles12.98 GiB1.55%1.519.47 GiB30.48 KiB
Wikipedia (en)6.38 GiB1.53%3.019.13 GiB1.11 KiB
DM Mathematics7.75 GiB1.24%2.015.49 GiB8.00 KiB
Ubuntu IRC5.52 GiB0.88%2.011.03 GiB545.48 KiB
BookCorpus26.30 GiB0.75%1.59.45 GiB369.87 KiB
EuroParl4.59 GiB0.73%2.09.17 GiB68.87 KiB
HackerNews3.90 GiB0.62%2.07.80 GiB4.92 KiB
YoutubeSubtitles3.73 GiB0.60%2.07.47 GiB22.55 KiB
PhilPapers2.38 GiB0.38%2.04.76 GiB73.37 KiB
NIH ExPorter1.89 GiB0.30%2.03.79 GiB2.11 KiB
Enron Emails0.88 GiB0.14%2.01.76 GiB1.78 KiB
Total1254.20 GiB5.91 KiB

(Epochs refers to the number of epochs elapsed after 1.2TB)

Usage

Install:

pip install -e .

To replicate pile

python the_pile/pile.py --interleave_output 30 --using pile_reprod

Use the pass 2 script here to complete shuffling.

Other

To force download all data:

python the_pile/pile.py --force_download

To generate fasttext training data for CC filtering (OWT2 only):

sudo apt install build-essential
python the_pile/pile.py --using owt2 --make_fasttext 

Manual Download Components

The following components need manual downloading. Either download them or comment out from pile.py.

Workflow

To propose a new dataset be added to the Pile, open an issue. Your issue should include a description of the dataset, its size, what language(s) it is in, a link to the data, and any other relevant information. If a project manger approves your proposal, they will change its label to Datasets and add it to Project: Datasets. Datasets that we elect to not include in the current version of the Pile will receive a Deferred or Declined label. While we welcome multilingual datasets and plan on including non-English datasets in the future, the initial release of the Pile will be English-only and all submissions of non-English datasets will be deferred.

To claim responsibility for implementing an unclaimed dataset, leave a comment on one of our unassigned issues. Once an dataset has been assigned to you, make the necessary changes to datsets.py and pile.py in a fork and submit a pull request. If you require, you can also submit a script for processing the data as shown here.

To raise an issue that is not proposing a new dataset, open an issue with the tag Feature Request or Bug as appropriate.

Data ready for final implementation should meet the following criteria:

In preparation for the initial release, we are no longer accepting additions to the master branch. If you would like to contribute a dataset, please submit the pull request to the Version2 branch.