Academic AI tools
Generative AI in library databases
Introduction
A growing number of library-subscribed databases now include generative AI features, going beyond other types of AI that have long been used in library resources, i.e. natural language processing, search optimization, predictive AI (recommending related sources).
It is crucial to understand that all generative AI tools, including those in library-subscribed resources, can generate misinformation (e.g. inaccurate summaries, imaginary sources). All AI-generated output must be carefully evaluated and verified.
Common types of generative AI integrations in library databases
Document summary
Generates a summary of an individual document/article, useful for determining relevancy and whether to read the full document.
Example: HeinOnline Article Summaries
Document-level chat
Provides a chat interface where users can interact with a single article - asking specific questions, clarifying concepts, etc.
Example: JSTOR AI Research Tool
Topic summary
Generates a summary of a topic or answers a question by synthesizing information from multiple articles in the database.
Example: O'Reilly for Higher Education Answers
Topic-level chat
Provides a chat interface where users can interact with multiple relevant articles – surfacing trends in the literature, comparing methodologies, etc.
Example: Oxford Academic AI Discovery Assistant
Library oversight
As generative AI features become more common in library-subscribed databases, the Library is reviewing these changes and advocating with vendors for transparency, privacy protections, risk mitigation, and usefulness for scholarly work. Vendor practices vary: some turn on AI features automatically, while others allow the Library to enable or disable them. To guide these conversations and decisions, the Library has identified expectations for vendors. When the Library has a choice about whether to activate AI features, it uses these same expectations to inform its decision.
Database vendors must:
- Clearly state that no user data will be shared or used to train AI models
- Clearly label AI generated output and provide citations to source material
- Disclose underlying models used, or, if models internally developed disclose training data
- Disclose defined set of data sources for generated output (e.g. database content only, or database content plus web sources)
- Must not incur any extra cost for the Library
- Must not solicit users to pay for extra features from within the Library-subscribed ecosystem
Database vendors should:
- Explain how output, recommendations or rankings are generated
- Have statements describing mitigation efforts re: bias, errors, environmental impacts, etc.
- Allow user choice in seeing summaries or engaging with chatbots, i.e. user can turn them off
The Library will continue to monitor developments in our subscribed resources and modify these expectations as needed.