How to Send Large Research Datasets to China

For research collaborations, the source dataset may already be on a public repository, cloud drive, server or object store. DataCrosslane can retrieve the required data from that source and deliver it to a supported destination for a team in China.

Start with the dataset source

Useful source information can be a Hugging Face or Zenodo page, a public project URL, Google Drive or OneDrive share, object-storage path, SFTP/WebDAV location, or direct HTTP/HTTPS download link.

A project page may be enough

If you do not have the exact file URLs, provide the dataset or project page. DOI, accession number and estimated size are helpful when available, but they are not required for every public dataset.

Transfer the whole dataset or only part of it

You can request a complete dataset, selected directories or specific file types. For example, an audio corpus can be limited to WAV, RTTM and JSON files while keeping the original directory hierarchy.

Large collections with many files

For multi-level datasets, you do not need to list every object individually. The root source and a clear description of the required subset are usually more useful.

Delivery for the research team

Baidu Netdisk and other supported storage destinations can be used depending on the recipient's workflow.