If you’re wondering how to prepare your research data for use in AI models, we have wonderful news for you. Good, old-fashioned best practices for sharing your data with humans also help prepare it for use by AI.

What kinds of data do AI models need?
Exact requirements are hard to come by for a few reasons: AI models vary, they’re updated frequently, and some are closely guarded secrets. But no matter who (or what) is using it, sharing data that is clean, consistent, and labeled helps ensure it’s interpreted correctly.
Start with these fundamentals
- Use consistent file names, terms, codes, and labels.
- Instead of color codes, use a categorical variable because that’s easier to compute.
- Contain your notes to designated places. Don’t sprinkle notes in with other data.
- Record nulls and NAs instead of leaving them blank.
For more info, see our quick guide to data cleaning and file naming convention template or contact the Illinois Research Data Service to set up a consultation.
What if I don’t want AI using my data?
AI comes with environmental and social costs, and there are lots of good reasons to approach it with caution.
There’s no standard way to express your AI preferences yet, but discussions are underway, such as the Creative Commons Signals project and this vocabulary from the Internet Engineering Task Force.
Whether you’re sharing research data with people or with AI,
clean, consistent, labeled data is a great place to start.
All Data Nudges are made with 100% human intelligence.
All Data Nudges are licensed under CC BY 4.0. You are free to share, adopt, or adapt them and cite the Illinois Research Data Service.