# Best practice for python package dependency?

**URL:** <https://discuss.dvc.org/t/best-practice-for-python-package-dependency/358>\
**Category:** Questions\
**Created:** [April 23, 2020, 11:26am UTC](https://discuss.dvc.org/t/best-practice-for-python-package-dependency/358 "2020-04-23T11:26:18Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![rabefabi](https://avatars.discourse-cdn.com/v4/letter/r/74df32/32.png) [@rabefabi](https://discuss.dvc.org/u/rabefabi)\
**Post date:** [April 23, 2020, 11:26am UTC](https://discuss.dvc.org/t/best-practice-for-python-package-dependency/358/1 "2020-04-23T11:26:19Z")

</div>

Hi all,

I’m figuring out how to structure my project and found your tutorials a very helpful introduction.

In them, however, you only rely on a single script for each stage, for example:

```auto
dvc run -d code/xml_to_tsv.py -d data/Posts.xml -o data/Posts.tsv \
          -f prepare.dvc \
          python code/xml_to_tsv.py data/Posts.xml data/Posts.tsv

```

[Source](https://dvc.org/doc/tutorials/pipelines)

Now my understanding is the following:  
If inside of `xml_to_tsv.py` are imports of some other of my libraries (e.g., a `my_special_xml_importer.py`), and I make changes to `my_special_xml_importer.py`, these changes would not be picked up by dvc, since `my_special_xml_importer.py` is not an explicit dependency of the stage, correct?

What’s the best practice here for bigger projects, where each DVC stage is not just contained in a single script?

Our use case will be the following: For each stage we’ll be having a jupyter notebook, which will import some of our python packages. I’m assuming I should create a stage like this:

```auto
dvc run -d my_notebook.ipynb -d code/my_lib.py -d data/Posts.xml -o data/Posts.tsv
  -f prepare.dvc
   papermill my_notebook.ipynb my_notebook_out.ipynb

```

Is this a good way, are there other ways, how are people with bigger projects dealing with this issue?

Thanks in advance,  
Fabi

---

<div class="post-metadata">

**Author:** ![shcheklein](https://yyz1.discourse-cdn.com/flex035/user_avatar/discuss.dvc.org/shcheklein/32/173_2.png) [@shcheklein](https://discuss.dvc.org/u/shcheklein)\
**Post date:** [April 24, 2020, 12:33am UTC](https://discuss.dvc.org/t/best-practice-for-python-package-dependency/358/2 "2020-04-24T00:33:36Z")

</div>

Hi @rabefabi!

The reasonable way is to put the code that is very specific to a stage into a separate directory:

`dvc run -d train ... train/train.py ...`

In this case, also use `.dvcignore` to exclude **pycache**.

I don’t know an easy way for DVC to analyze the actual graph of dependencies build and maintain them. It can be probably done with some custom scripts though.

---

<div class="post-metadata">

**Author:** ![rabefabi](https://avatars.discourse-cdn.com/v4/letter/r/74df32/32.png) [@rabefabi](https://discuss.dvc.org/u/rabefabi)\
**Post date:** [April 27, 2020, 6:52am UTC](https://discuss.dvc.org/t/best-practice-for-python-package-dependency/358/3 "2020-04-27T06:52:37Z")

</div>

Putting the notebooks in each stages subdirectory should work for us, thanks for the hint
