NIyueeE commited on
Commit
83e84f8
·
verified ·
1 Parent(s): d93974e

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +2 -14
README.md CHANGED
@@ -39,7 +39,7 @@ NVIDIA A100-SXM4-40GB. Num GPUs = 1. Max memory: 39.494 GB.
39
  Torch: 2.10.0+cu128. CUDA: 8.0. CUDA Toolkit: 12.8. Triton: 3.6.0
40
  Bfloat16 = TRUE. FA [Xformers = 0.0.35. FA2 = False]
41
 
42
- Num examples = 1,227 | Num Epochs = 5 | Total steps = 50
43
  Batch size per device = 128 | Gradient accumulation steps = 1
44
  Total batch size (128 x 1 x 1) = 128
45
  Trainable parameters = 13,181,952 of 866,167,872 (1.52% trained)
@@ -65,22 +65,10 @@ See [`finetune_cocreator_coclab.ipynb`](./finetune_cocreator_coclab.ipynb) for t
65
  | Optimizer | adamw_8bit |
66
  | Learning rate | 5e-5 (cosine schedule) |
67
  | Max steps | 50 |
68
- | Epochs | 5 |
69
  | Gradient checkpointing | unsloth |
70
  | Resolution | 800×450 (resized) |
71
 
72
- ## Training Results
73
-
74
- | Step | Loss |
75
- |------|------|
76
- | 10 | 20.36 |
77
- | 20 | 12.69 |
78
- | 30 | 8.22 |
79
- | 40 | 7.06 |
80
- | 50 | 6.79 |
81
-
82
- The model was trained for 50 steps (~5 epochs), with loss dropping from 20.36 to 6.79, indicating steady convergence on the driving scene causal understanding task.
83
-
84
  ## Usage
85
 
86
  ```python
 
39
  Torch: 2.10.0+cu128. CUDA: 8.0. CUDA Toolkit: 12.8. Triton: 3.6.0
40
  Bfloat16 = TRUE. FA [Xformers = 0.0.35. FA2 = False]
41
 
42
+ Num examples = 1,227 | Num Epochs = 7 | Total steps = 50
43
  Batch size per device = 128 | Gradient accumulation steps = 1
44
  Total batch size (128 x 1 x 1) = 128
45
  Trainable parameters = 13,181,952 of 866,167,872 (1.52% trained)
 
65
  | Optimizer | adamw_8bit |
66
  | Learning rate | 5e-5 (cosine schedule) |
67
  | Max steps | 50 |
68
+ | Epochs | 7 |
69
  | Gradient checkpointing | unsloth |
70
  | Resolution | 800×450 (resized) |
71
 
 
 
 
 
 
 
 
 
 
 
 
 
72
  ## Usage
73
 
74
  ```python