# Large dataset, dvc pull/add/push jobs options

**URL:** <https://discuss.dvc.org/t/large-dataset-dvc-pull-add-push-jobs-options/1496>\
**Category:** Questions\
**Created:** [February 3, 2023, 2:36am UTC](https://discuss.dvc.org/t/large-dataset-dvc-pull-add-push-jobs-options/1496 "2023-02-03T02:36:43Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![dsa934](https://avatars.discourse-cdn.com/v4/letter/d/ebca7d/32.png) [@dsa934](https://discuss.dvc.org/u/dsa934)\
**Post date:** [February 3, 2023, 2:36am UTC](https://discuss.dvc.org/t/large-dataset-dvc-pull-add-push-jobs-options/1496/1 "2023-02-03T02:36:43Z")

</div>

Hello , DVC users

I am a newbie who wants to apply dvc to large dataset management.  
To understand dvc, I did the following experiment.

```auto
## normal transfer case 

# In DVC_Main directory 
git init
dvc init
dvc remote add -d <storage_name> <local_data_storage_url>

dvc add mesh_dataset # mesh_dataset (40GB)
dvc push 

# In DVC_Sub directory 
git init
dvc init

#copy & pasted files from ( dvc_main directory) -> .dvc/config , .mesh_dataset.dvc files 
dvc pull 

```

The above process is a code that tests whether the data worked in the main directory can be loaded from the sub directory.  
But it took too long (about 80 minutes)

So I looked for options and there was a -j option for dvc add/pull/push.

```auto
## parallel transfer case 

# In DVC_Main directory 
git init
dvc init
dvc remote add -d <storage_name> <local_data_storage_url>
dvc remote modify jobs 64

dvc add -j 64 --to-remote mesh_dataset # mesh_dataset (40GB)

# In DVC_Sub directory 
git init
dvc init

#copy & pasted files from ( dvc_main directory) -> .dvc/config , .mesh_dataset.dvc files 
dvc pull -j 64

```

The parallel transmission method should be faster than the normal transmission method, but I don’t understand why there is no speed difference.

In fact, the parallel transmission method is slightly faster, but it was a difference that appeared because the ‘dvc push’ process was omitted in the dvc\_main directory.

![image](https://canada1.discourse-cdn.com/flex035/uploads/dataversioncontrol/original/1X/29914ea01d5fc60bbbfb026c5970afc795dc9522.png)

According to the explanation above, I have 1 cpu (8 cores, 14 logical cores), so the default value = 34.

Can you explain why there is no difference in transfer speed between jobs option 32 and 64?

Or does dvc not support parallel transmission of large data?

---

<div class="post-metadata">

**Author:** ![ronan](https://avatars.discourse-cdn.com/v4/letter/r/3ec8ea/32.png) [@ronan](https://discuss.dvc.org/u/ronan)\
**Post date:** [February 7, 2023, 12:50pm UTC](https://discuss.dvc.org/t/large-dataset-dvc-pull-add-push-jobs-options/1496/2 "2023-02-07T12:50:47Z")

</div>

Hi @dsa934 !  
`dvc` does push and pull data in parallel, whether you use the `-j` option or not. However, parallelisation has diminishing returns due to resource constraints (network bandwidth, unparallelisable per-job overhead, …) From your testing, I guess that 32 is already past the threshold where adding more jobs doesn’t bring any speed-up.

---

<div class="post-metadata">

**Author:** ![dsa934](https://avatars.discourse-cdn.com/v4/letter/d/ebca7d/32.png) [@dsa934](https://discuss.dvc.org/u/dsa934)\
**Post date:** [February 7, 2023, 2:57pm UTC](https://discuss.dvc.org/t/large-dataset-dvc-pull-add-push-jobs-options/1496/3 "2023-02-07T14:57:47Z")

</div>

Hi @ronan !

According to you basically 32 means threshold (because 1 cpu has 8 cores, formula : 4 \* cpu\_count() ), does that mean we need to increase the number of actual cpus to improve speed?

However, when I experimented with smaller numbers, such as 12 or 16 instead of 64, there was no change in speed.
