# Multiple overlapping datasets

**URL:** <https://discuss.dvc.org/t/multiple-overlapping-datasets/576>\
**Category:** Questions\
**Created:** [December 7, 2020, 2:31pm UTC](https://discuss.dvc.org/t/multiple-overlapping-datasets/576 "2020-12-07T14:31:19Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![kazimpal](https://avatars.discourse-cdn.com/v4/letter/k/d07c76/32.png) [@kazimpal](https://discuss.dvc.org/u/kazimpal)\
**Post date:** [December 7, 2020, 2:31pm UTC](https://discuss.dvc.org/t/multiple-overlapping-datasets/576/1 "2020-12-07T14:31:19Z")

</div>

Hi,

I’m considering using dvc for a project and am wondering if it fits the following use case. Say I have N images which form a dataset, but for any given experiment I may only want to train on some subset of the images. What I would like is to be able to determine which images in the subset are missing locally and only pull those images from the file store. Ideally these subsets could also have names and I’d be able to just do dvc pull datasubset3. Is something like this possible?

Thanks

---

<div class="post-metadata">

**Author:** ![jorgeorpinel](https://yyz1.discourse-cdn.com/flex035/user_avatar/discuss.dvc.org/jorgeorpinel/32/46_2.png) [@jorgeorpinel](https://discuss.dvc.org/u/jorgeorpinel)\
**Post date:** [December 7, 2020, 6:59pm UTC](https://discuss.dvc.org/t/multiple-overlapping-datasets/576/2 "2020-12-07T18:59:43Z")

</div>

Hello,

> [@kazimpal](#):
>
> Say I have N images which form a dataset

OK so let’s assume that dataset is a [directory which is tracked](https://dvc.org/doc/command-reference/add#adding-entire-directories) by DVC.

> [@kazimpal](#):
>
> What I would like is to be able to determine which images in the subset are missing locally and only pull those images

Once you have determined which files you can pull those specifically, yes. But it has to be done one by one (you can print the entire list to a file and then use a shell script to `dvc pull` each name). This is what we call “granularity support” in most of our commands, including `push` and `pull`.

> [@kazimpal](#):
>
> Ideally these subsets could also have names and I’d be able to just do dvc pull datasubset3

DVC doesn’t have such a feature, but we’re open to feature requests in [GitHub - iterative/dvc: 🦉 ML Experiments Management with Git](http://github.com/iterative/dvc).

Note that we are also about to release a wildcard feature called “globbing” which you can see here [pull: add glob option by ju0gri · Pull Request #5032 · iterative/dvc · GitHub](https://github.com/iterative/dvc/pull/5032), maybe that will be enough for your case? Depending on your dataset naming/file structure.

Best
