File size: 521 Bytes
d9cf33b d2eb966 d9cf33b 3cb4a20 d9cf33b 99a750a d2eb966 99a750a d2eb966 f5110dc 99a750a |
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 |
# Train
## Tokenizer
```bash
cd scripts
python -m venv venv
source venv/bin/activate
pip install -U -r requirements.in
```
```bash
python -B train_tokenizer.py
```
## Dataset
```bash
cd scripts
python -m venv venv-lit
source venv-lit/bin/activate
pip install -U -r requirements-lit.in
```
```bash
python -B prepare_pretrain_dataset.py
```
## Model
```bash
cd scripts
python -m venv venv-lit
source venv-lit/bin/activate
pip install -U -r requirements-lit.in
```
```bash
litgpt pretrain --config ./model.yaml
```
|