TL;DR:
Should i use git submodules as a way to include large data in a repository so that the main repository does not bloat?
i am using git to keep track of configurations,settings and scripts for analysis of weather stations i manage for my university. Each Station has its own repository. On top of that i am using gitlab to host the remotes and keep track of issues with these stations. I really like this workflow.
Now theses stations produce a lot of data in different formats. Some generate just a few MB/month and some, in rare cases, 4GB/month running for multiple years.
I already manage the data with a database(influxdb and grafana) where the data is simplified a little bit for monitoring and quickliy visualising data. But i really want to keep the RAW original Data somewhere, in case of some issue with the database and just as a dumb backup. Mostly these are csv, json or some other textbased files. And i want that backup to be redundant, reliable and accessible. (Longterm the data will go into some kind of AWS glacier but thats another story)
Obviously, i dont want to keep it in the main repo for the station since this will bloat the repo and introduce problems on gitlab. But i also like the thought that the data is connected somehow to the station it belongs to.
So how do i store the data without bloating my repo?
1. Just store it on a file server and document it then..
So this one is kind of lame and does not feel complicated enough :D. No seriously since the number of stations will grow and i sometimes want to just do something like git clone --recurse-submodules ... and have everything i need for that station, i want the data somehow to be related to the repo.
2. Thats what git lfs is made for..
I think this would be the most obvious solution, but i have one problem with that. I would need to maintain my own instance of gitlab on a server that has enough storage since the limit on gitlab.com is 10GB. Maybe i will manage my own instance in the future where this restriction does not exist but i think i want a dedicated local server for that. But i dont want to manage my own raid and everything that comes with reliable backup storage. That i want to outsource to some storageprovider or the datacenter at our university. In the gitlab docs is a section about hosting lfs externally but it seems like this external storage needs to support lfs as well.? And it does not seem like i just need to install git lfs there... So maybe still possible, but i guess setting up the backend for that is more complicated.
3. Using submodules like a fool
My idea was to just use an extra repository for the data. This repo would be hosted on a cheap dedicated fileserver with enough space and ssh connection. I then would add this remote as a submodule to the mainrepo. I tested this already and it works but:
a. Adding files is slow
i expected this, its ok
b. compressing when pushing and cloning is slow
no way around that i guess..
c. writing and receiving is slow
Writing objects: 100% (78/78), 2.28 GiB | 4.24 MiB/s, done.
Receiving objects: 25% (24/95), 153.55 MiB | 4.56 MiB/s
Why? Isnt this just a ssh connection? It tops out at around 4MiB/s while i know that my connection is a lot faster. Around 100MiB/s when using scp to the same server in both directions
I kind of like this solution the most since it does require the least amount of additional software to get it working.
May i get punished by the git gods.
Is it possible to optimize the submodule workflow or should i do it differntly since this just is not what git is made for? I guess there is a very straight forward way doing this which i am not seeing because i am in tunnel vision right now.