mirror of
https://huggingface.co/datasets/QuixiAI/open-instruct-uncensored
synced 2026-09-11 11:07:35 +02:00
No description
- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .gitattributes | ||
| open-instruct-uncensored.jsonl | ||
| open-instruct.jsonl | ||
| README.md | ||
| remove_refusals.py | ||
| license |
|---|
| apache-2.0 |
This is Allen AI's open-instruct dataset.
It is used to train the Tulu family of models.
- https://huggingface.co/allenai/tulu-7b
- https://huggingface.co/allenai/tulu-13b
- https://huggingface.co/allenai/tulu-30b
- https://huggingface.co/allenai/tulu-65b
I have done the following:
- Download the open-instruct repo
- Execute the scripts/prepare_train_data.sh modified to download the "unfiltered" version of sharegpt dataset
- Merged data/processed/**/*.jsonl into a single "open-instruct.jsonl"
- Executed my "remove_refusals.py" against that "open-instruct.jsonl" to produce a "open-instruct-uncensored.jsonl"
I am currently training this "open-instruct-uncensored.jsonl" to a new model series named ehartford/tulu-uncensored
More info to come.