Scholarly source landscape
Scholarly source landscape
Survey date: 2026-08-30. This is a product-compatibility review of published provider terms, not legal advice. Terms and API behavior can change; re-check the cited first-party pages when cutting an implementation brief.
Summary
Three of the 22 shipped integrations are incompatible as currently implemented: IEEE Xplore, INSPIRE-HEP, and Springer Nature. Part A also finds 4 compatible sources, 4 compatible sources with an apparently unmet obligation, and 11 unresolved sources whose published terms do not establish permission for at least one required act. Unresolved does not mean compatible.
Part B surveys 27 candidates: 6 ADOPT, 4 ADOPT-WITH-KEY, and 17 REJECT. The survey is deliberately metadata-first. A license for metadata never silently extends to linked articles, files, images, or other content.
Decision rule
The terms verdict asks whether resmon may (1) keep the retrieved record indefinitely in its local SQLite database, (2) copy that private database to user-controlled cloud storage, (3) do so in a freely installable MIT-licensed application that redistributes no records, and (4) retain the first-retrieved record without a provider-mandated refresh. “Compatible with an obligation” means the four acts are permitted only if the named attribution, notice, contact, or license condition is met. “Unresolved” means the published terms do not answer the question; public access alone is not treated as permission.
Part A compares the terms with the clients at
upstream/main commit 46f9394.
Part B uses ADOPT only when no key is needed, ADOPT-WITH-KEY for a free
key or equivalent free registration, and REJECT for paid/institutional
access, an incompatible or missing reuse grant, or an API that cannot perform
the monitoring job.
Part A — shipped-source terms re-check
The four numbered results in each row correspond to the four questions above. Where a result is field-scoped, the field boundary is part of the verdict.
| Shipped source | Verdict | Four-question result, governing clause, and current obligation |
|---|---|---|
| arXiv | COMPATIBLE WITH AN APPARENTLY UNMET ACKNOWLEDGMENT | 1 yes; 2 yes; 3 yes; 4 yes, for descriptive metadata only. The API terms define titles, abstracts, authors, identifiers, and classifications as CC0 metadata and expressly permit retrieving, storing, transforming, and sharing it; e-prints remain separately licensed. The API access page requests the sentence “Thank you to arXiv for use of its open access interoperability” and bars names/branding that imply endorsement. The requested sentence was not found in the reviewed tree; ordinary source identification does not itself imply endorsement. |
| bioRxiv | UNRESOLVED | 1 unknown; 2 unknown; 3 record-dependent; 4 no refresh duty found. bioRxiv provides free machine access and API metadata under its TDM guidance, but that page does not grant blanket indefinite or cloud-copy rights. The API returns abstracts and item-level licenses, while bioRxiv’s author guidance permits choices ranging from CC0/CC licenses to “no reuse.” The current client retains abstracts without enforcing the returned license, so API availability does not resolve durable copying. |
| CORE | UNRESOLVED; MANDATORY CONTACT/LICENSE AND ATTRIBUTION APPARENTLY UNMET | 1–4 unresolved until the required license is obtained and reviewed. CORE’s terms reserve database rights and require a license for API/non-ODC dataset use; its FAQ requires developers of public search/discovery products to contact CORE, send a use summary, and show its “powered by CORE” snippet. The API page makes access subject to those terms. No contact/license or required snippet was found in the reviewed tree, the published pages do not establish that the future license permits all four acts, and a CORE license would not automatically convey rights in third-party abstracts. |
| Crossref | UNRESOLVED FOR THE CURRENT NORMALIZED RECORD | 1 yes; 2 yes; 3 yes; 4 yes for factual metadata and Crossref-created data, but 1–3 unknown for abstracts. Crossref says almost all bibliographic metadata is reusable and its own data is CC0, but expressly leaves abstract copyright with publishers/authors in Metadata Retrieval. No provider-mandated refresh was found for otherwise reusable fields, and Crossref recommends caching in its REST access guidance. The current client stores abstracts without an item-level rights gate. |
| DataCite | COMPATIBLE | 1 yes; 2 yes; 3 yes; 4 yes. The Data File Use Policy applies CC0 to DataCite DOIs and deposited metadata with no use conditions. Linked resources and third-party privacy/publicity rights remain outside that grant; attribution and non-misleading modification are requested norms rather than license conditions. |
| DBLP | COMPATIBLE | 1 yes; 2 yes; 3 yes; 4 yes. DBLP’s license FAQ releases all DBLP metadata under CC0 for copying, transformation, building upon, and commercial or noncommercial use without permission; linking back is appreciated, not required. Its search API documentation repeats the CC0 metadata boundary. Linked publications are not DBLP metadata. |
| DOAJ | COMPATIBLE WITH AN APPARENTLY UNMET NAME-USE OBLIGATION | 1 yes; 2 yes; 3 yes; 4 yes for API metadata. DOAJ’s terms, version 1.6 place journal/article API metadata under CC0, but also say the “Directory of Open Access Journals” name and “DOAJ” acronym may not be reproduced without prior written permission. The FAQ confirms that anyone may ingest the metadata without attribution and that article rights remain separate. Resmon displays “DOAJ”; no written permission was found in the reviewed tree or repository records. |
| ERIC | UNRESOLVED | 1 unknown; 2 unknown; 3 encouraged but not licensed for every stored field; 4 no refresh duty found. The ERIC OpenAPI specification and API page encourage external applications and identify description as an abstract that may come from ERIC, a publisher, or an author. IES describes citation metadata as free and open to use in its public-access policy, but publishes no blanket durable-copy grant for publisher/author abstracts. The current client stores those descriptions. |
| Europe PMC | UNRESOLVED FOR THE CURRENT NORMALIZED RECORD | 1–3 have no blanket permission across the abstract corpus; 4 no refresh duty found. EMBL-EBI’s terms expect attribution but leave original-owner restrictions in force. Europe PMC’s copyright page says API/OAI/bulk material remains subject to each article’s license, while its developer page encourages applications built from open content and metadata. An item-level open license can clear an individual abstract, but the current client stores abstractText across records without enforcing one. |
| HAL | UNRESOLVED BECAUSE FIRST-PARTY TERMS CONFLICT | HAL’s Principles and publishing workflow apply CC0 to metadata, which would make 1–4 yes. HAL’s current legal guidance and OAI documentation instead require source citation and say extracted data may not be used commercially. Because both statements remain published, 1–4 are not terms-cleared for the unrestricted product model until HAL resolves which governs the Solr research API. Deposited files remain separate under either reading. |
| IEEE Xplore | INCOMPATIBLE AS SHIPPED | 1 no; 2 no; 3 no outside a narrow institution-bound license; 4 no because termination requires deletion. The IEEE Xplore API Terms limit the grant to noncommercial educational/research/scientific activity within the licensee institution (§ Grant), prohibit a retrieval/indexing application (§4(c)) and bulk retention (§4(f)), and require removal of Content on termination (§12). Resmon’s persistent searchable database, Drive backup, public installer, and no-delete model conflict; IEEE’s supported use cases do not override those clauses. |
| INSPIRE-HEP | INCOMPATIBLE AS SHIPPED | 1–4 yes for CC0 fields and only qualifying abstracts; no for the record currently stored without provenance enforcement. INSPIRE’s Terms §5 allow abstract reuse only when abstracts.source is arXiv or CERN, restrict keyword-field sources, and forbid email download. The current client requests abstracts and stores the first one without checking abstracts.source; that field-level license violation makes the normalized record incompatible. |
| medRxiv | UNRESOLVED | 1 unknown; 2 unknown; 3 record-dependent; 4 no refresh duty found. medRxiv’s TDM page makes metadata available through APIs but does not grant blanket indefinite retention or cloud backup; the FAQ describes article licenses ranging from CC0/CC BY to no reuse, and the API returns the item license alongside the abstract. The shared current client stores abstracts without enforcing that license. |
| NASA ADS | UNRESOLVED; CONTACT REQUEST APPARENTLY UNMET | 1 only for personal use and subject to contact for regular storage; 2 unknown; 3 uncertain; 4 no refresh duty found. ADS Terms §§2–3 forbid API-data redistribution, ask users who regularly download/store ADS data to contact ADS first, and prohibit systematic partial/full corpus copies; abstracts and full text remain publisher property. No prior contact was found, and the terms do not say whether a private Google Drive service-provider copy is prohibited third-party redistribution. |
| OpenAIRE | COMPATIBLE WITH APPARENTLY UNMET CC BY OBLIGATIONS | 1 yes; 2 yes; 3 yes; 4 yes for Graph metadata. OpenAIRE’s API terms and Graph license allow commercial and noncommercial reuse under CC BY with OpenAIRE acknowledgment. CC BY 4.0 requires attribution, a license link, and an indication when changes were made. Those elements were not found together in the reviewed record-display/docs surfaces; linked full text remains separate. |
| OpenAlex | COMPATIBLE | 1 yes; 2 yes; 3 yes; 4 yes for OpenAlex data. OpenAlex says all of its data, however accessed, is CC0 in How it’s built and its access overview. The broad unauthorized-copy language in the 2026 Terms of Service is read together with that explicit CC0 authorization; linked works remain outside the grant. |
| OSTI.GOV | UNRESOLVED FOR THE CURRENT NORMALIZED RECORD | 1 yes for saving bibliographic metadata; 2 unresolved for potentially copyrighted descriptions; 3 field-dependent; 4 no refresh duty found. OSTI’s FAQ permits saving bibliographic metadata and identifies the API description field as metadata. Its site policies clearly permit reuse of public-domain government material, ask for OSTI credit, and identify an OpenAIRE-derived CC BY subset, but do not grant blanket private-cloud copying rights across every description. The API docs define the delivered fields. Narrow to public-domain/licensed records or obtain clarification before treating the current record as cleared. |
| PLOS | COMPATIBLE WITH AN APPARENTLY UNMET DISPLAY OBLIGATION | 1 yes; 2 yes; 3 yes; 4 yes for PLOS API material under its applicable open license. PLOS’s Terms of Use make articles and accompanying material generally CC BY or comparably permissive, subject to source citation. Its API Display Policy requires the exact nearby phrase “Data Provided by PLOS,” and its TDM guidance welcomes desktop/mobile/web applications. The required phrase was not found in the reviewed tree. |
| PubMed / NCBI E-utilities | UNRESOLVED; E-UTILITIES IDENTITY/NOTICE OBLIGATIONS APPARENTLY UNMET | 1–2 unknown for records containing abstracts; 3 not established for all abstracts; 4 no applicable provider-mandated refresh found. NCBI’s policies warn that database material can retain third-party copyright, require its disclaimer/copyright notice to be evident, and require distributed software to send tool and email. NLM’s separate FTP download terms also warn that PubMed abstracts may be publisher/author copyrighted, but their courtesy/currentness clauses are not treated here as E-utilities obligations without a first-party bridge. The current client stores abstracts and does not send the E-utilities identity parameters; the applicable policy notices were not found. |
| Semantic Scholar | UNRESOLVED; ATTRIBUTION OBLIGATION | 1–2 contemplated but not granted across the corpus; 3 only under accompanying licenses; 4 no refresh duty found. The Semantic Scholar API License §§2 and 4 recognizes that data may be accessed, stored, or transmitted, but makes S2 data subject to accompanying licenses such as CC BY-NC or ODC-BY and preserves third-party content terms; it also requires “Semantic Scholar” attribution. The API overview provides no blanket title/abstract grant, and the Graph response does not supply a general per-record license gate. The current client stores abstracts without capturing or surfacing an applicable license. |
| Springer Nature Meta API | INCOMPATIBLE AS SHIPPED | 1 no; 2 no; 3 no for non-OA abstracts beyond personal use; 4 not curative. Springer Nature’s Additional API Terms, “Use of abstracts” call abstracts copyrighted material, grant only a limited personal-use license, and require express consent for other reuse; institutional TDM storage is generally project-duration-only and confined to an internal server. The Meta v2 documentation confirms that the API returns abstracts. The current client stores every returned abstract without an OA-status/license gate. |
| Zenodo | COMPATIBLE | 1 yes; 2 yes; 3 yes; 4 yes for metadata other than email addresses after permitted access. Zenodo Terms §7, policies, and developer documentation apply CC0 to titles, authors, descriptions, and other metadata, allow harvesting, and exclude bulk email-address download. The Terms separately restrict service access to non-military purposes; the CC0 waiver permits unrestricted reuse of metadata lawfully obtained. Deposited files and full text follow their own licenses. |
Part A disposition
The three incompatible integrations are live product issues, not documentation niceties:
- IEEE Xplore needs a separate written license or removal/disablement of the source; its published terms conflict with the product model.
- INSPIRE-HEP needs field-level provenance enforcement before an abstract is retained.
- Springer Nature needs an OA/license gate that excludes non-permitted abstracts, express permission, or removal/disablement of the source.
The 11 unresolved integrations—bioRxiv, CORE, Crossref, ERIC, Europe PMC, HAL, medRxiv, NASA ADS, OSTI, PubMed, and Semantic Scholar—should not be represented as terms-cleared until the stored field set is narrowed, first-party conflicts are resolved, or the provider/item-level license establishes all four uses.
Part B — candidate survey
Rows are intentionally brief-shaped: each contains the proposed slug and gap, endpoint/auth, rate, keyword semantics, date behavior, pagination, the four-question terms result, and a decision/priority. “Undocumented” means the cited first-party material was checked and did not answer the point.
| Candidate / proposed slug / coverage | Endpoint, authentication, and rate | Keyword, date, and pagination contract | Four-question terms result | Decision |
|---|---|---|---|---|
ClinicalTrials.gov API v2 / clinicaltrials — worldwide interventional and observational trial registrations/results; the strongest clinical-trials gap fill (official data/API overview). |
GET https://clinicaltrials.gov/api/v2/studies; auth: none, and access is free under the site terms. Rate: none published in those cited pages. |
Bare words undergo documented relaxation; emit explicit Boolean AND (also OR/NOT) under the query guide. Date: inclusive AREA[...]RANGE[YYYY-MM-DD,YYYY-MM-DD], day granularity. Pagination: pageSize default 10/max 1,000 plus pageToken/nextPageToken; no total ceiling is documented in the migration guide. |
Q1 yes; Q2 undocumented; Q3 yes; Q4 yes under the no-redistribution design. The terms address retention and impose attribution, currentness, processing-date, modification, and no-proprietary-rights duties when data are published or distributed, but do not say whether a private service-provider cloud copy is permitted. Public availability is not treated as an answer to Q2. | REJECT FOR NOW. Google Drive backup is part of the product, so adopt only after provider clarification or an enforceable source-specific backup exclusion. |
GovInfo Search Service / govinfo — official U.S. legislative, judicial, regulatory, standards-adjacent, hearing, and technical-report material from all three federal branches (API data scope). |
POST https://api.govinfo.gov/search; auth: free API.data.gov key. Published defaults are 36,000/hour, 1,200/minute, and 40/second in the official API README. |
Spaces imply AND, with explicit Boolean and field operators in Search Operators. Date: publishdate:range(YYYY-MM-DD,YYYY-MM-DD), day granularity, under the search help. Pagination: pageSize/opaque offsetMark; published and collection endpoints document max 1,000/page and cursor traversal beyond 10,000, with no total ceiling stated in the API README. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes for GPO-created metadata and U.S.-government works, not embedded third-party material. GovInfo’s policies expressly support downloading metadata/content packages into local digital collections and explain that 17 U.S.C. §105 places U.S.-government works in the public domain, allowing copying to private storage without a refresh duty. The same clause warns that government publications can embed copyrighted material, so retain only official metadata/teasers and links. | IMPLEMENTED (Delegation 08, 2026-09-06). Official bibliographic fields only; no teasers or document text. Search-specific pagination, including the response cursor, verified live with DEMO_KEY. |
Dryad / dryad — curated research datasets across disciplines (reuse guide). |
GET https://datadryad.org/api/v2/search; auth: none for public search. Rate: none published numerically in the API/search guidance. |
Search covers most metadata; quotes, suffix wildcard, and NOT/- are documented, but bare multi-term combination is undocumented in the search guide. Date: publishedSince/publishedBefore in the API reference, ISO date/time; finest documented input is a timestamp. Pagination: page/per_page, max 100/page; no total ceiling published in that reference. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes. Dryad’s Terms §§1 and 4 publish datasets under CC0 and encourage reuse; its reuse guide says reuse, modification, sharing, redistribution, and commercial use need no permission. Citation is strongly recommended, not legally required. | ADOPT — priority 1. Best direct dataset gap fill; store metadata only and preserve DOI/version context. |
NDL Search / ndl_search — Japanese national bibliography, periodical index, books, articles, and other Asian-language collections (API scope). |
GET https://ndlsearch.ndl.go.jp/api/sru; auth: none for the dpid=open PDM/CC0 slice, including for-profit use, under the provider table. Numeric rate not published; continuous users are asked to identify themselves and excessive/concurrent traffic may be blocked under the usage terms. |
CQL with explicit AND/OR; do not invent a bare-space rule. Date: from/until publication-year fields, year granularity, with an official example in the specification overview. Pagination: startRecord/maximumRecords, default 200/max 500, but only the first 500 hits are retrievable under the current v1.4 specification. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes only for the dpid=open PDM/CC0 provider slice. The provider table says that slice requires no application and distinguishes all other provider-specific conditions. The API terms require “NDL Search API” credit and any provider/CC BY credit; preserve provider/license fields and exclude non-open records. |
ADOPT — priority 1. Strong geography/language/books gap; the brief must enforce dpid=open and partition queries to stay below the 500-hit ceiling. |
Open Library Search API / openlibrary — global work/edition bibliographic metadata for books and humanities (Search API). |
GET https://openlibrary.org/search.json; auth: none. Rate: 1 request/second unidentified or 3/second with an app/contact User-Agent; regular clients should identify and cache under the API guidelines. |
q is a Solr query, but plain-space semantics are undocumented in the Search API; use explicit syntax. Date: exact/range publish_year and first_publish_year, year granularity, in the search guide. Pagination: offset/limit or 1-based page/limit; no total ceiling published. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes, with provenance retained. Using Open Library Data says catalog source data were donated to the public domain and new contributions use CC0; developer licensing cautions that pre-existing rights can survive in some sources/jurisdictions. Low-volume identification is an operational obligation. | ADOPT — priority 2. Low-friction books/humanities source with caching aligned to resmon’s no-refresh design. |
NIST Resource Metadata Management API / nist_rmm — NIST papers, patents, software, APIs, and data-resource metadata; useful for engineering/technical-report coverage (RMM API). |
GET https://data.nist.gov/rmm/papers; auth: none documented by the OpenAPI schema. Rate: none published. |
searchphrase exists but multi-term combination is undocumented. Date: lower-bound from_date=YYYY-MM-DD only, day granularity; no upper bound. Pagination: offset skip plus limit; defaults 0/10 and no maxima or total ceiling are published in the OpenAPI schema. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes for NIST-created metadata/works, with attribution. NIST’s copyright/licensing statement says ordinary NIST employee works are generally not U.S.-copyrighted, grants worldwide reuse where NIST may hold foreign rights, and conditions use on appropriate NIST acknowledgment; third-party and Standard Reference Data exceptions remain. | ADOPT — priority 2. Valuable engineering/government research slice; require NIST credit and a metadata-only/exception-aware field policy. |
OAPEN Library / oapen — multilingual peer-reviewed open-access books and chapters across humanities/social sciences (metadata page). |
GET https://library.oapen.org/rest/search?query=...; auth: none. Rate: none published in the cited REST/metadata pages. |
Official examples use explicit AND and fielded/quoted queries; bare-space semantics are undocumented in the REST guide. Date: Solr ranges on dc.date.issued_dt or accession dates, timestamp granularity. Pagination model and result ceiling are undocumented in that guide; the alternative OAI-PMH feed is resumable but has no keyword query. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes for metadata. OAPEN states that all feeds are CC0 and are intended for library/aggregator integration on its metadata page; book files follow item licenses. | IMPLEMENTED (Delegation 08, 2026-09-06). CC0 metadata only; whole publication years. Pagination resolved in the provider REST PDF: expand, limit, offset; distinct live pages verified. |
DPLA / dpla — U.S. library, archive, and museum metadata for humanities and cultural-history discovery (developer overview). |
GET https://api.dp.la/v2/items; auth: free API key obtained through DPLA’s key policy. DPLA publishes no routine numeric rate; it reserves intervention for abusive/degrading use in that policy. |
q searches all text fields; multiple terms default to AND and explicit Boolean is supported in the request docs. Date: sourceResource.date.after/before, accepting year or ISO-style values according to record precision. Pagination: 1-based page, page_size max 500, page max 100: a 50,000-result ceiling. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes for metadata only. DPLA Terms §5.2 dedicate Metadata to CC0; content objects and visual assets retain separate rights. | ADOPT-WITH-KEY — priority 3. Useful humanities breadth, but less scholarly-specific and ceiling-aware partitioning is required. |
Europeana Search API / europeana — multilingual European cultural-heritage metadata, especially manuscripts, newspapers, books, and non-Anglophone humanities (Search API). |
GET https://api.europeana.eu/record/v2/search.json; auth: free key under Accessing APIs. Europeana says read APIs are free and unthrottled but asks clients to leave a few milliseconds between calls in that FAQ. |
Lucene/Solr syntax; consecutive unquoted terms default to AND under the query syntax. Date: ISO 8601 ranges exist for Europeana timestamp_created/timestamp_update (ingestion/update, not original publication), but the finest accepted query precision is undocumented; provider publication fields are heterogeneous strings. Pagination: rows max 100, basic offset first 1,000, or opaque cursor/nextCursor through the complete set. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes for Europeana metadata. The API FAQ and licensing framework apply CC0 to metadata; linked digital objects follow edm:rights. |
ADOPT-WITH-KEY — priority 3. Strong multilingual humanities coverage, but date semantics must be labelled as record ingestion/update. |
OpenReview API v2 / openreview — conference submissions and public decision/review metadata, especially computer-science proceedings (Notes model). |
GET https://api2.openreview.net/notes/search; auth: free profile/credentials under the client setup. Rate: none published numerically; terms reserve discretionary upper limits. |
query/term/terms and prefix/terms/exact exist, but space combination is undocumented in the OpenAPI definition. Search has no upstream date parameter; notes separately expose millisecond cdate/mdate for local filtering in Note fields. Pagination: limit/offset; ceilings undocumented. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes for metadata only. Terms, Metadata Dedication apply CC0 to metadata; public comments/configuration are CC BY, and submissions/review text can have separate restrictions. | DECLINED for Delegation 08 (2026-09-06). Anonymous public search works, contrary to the keyed verdict. Live /notes/search rejects offset 10000 + limit 1 (10,000-result window), and rejects sort=pdate:desc. The search contract has no date-range filter; pdate is acceptance, not cdate. Local filtering cannot establish an exhaustive narrow window beyond the ceiling; the frozen outcome channel has no incomplete-window state. Metadata CC0 remains compatible; no terms prohibition is claimed. |
Gallica SRU / gallica — BnF/partner books, periodicals, manuscripts, maps, and humanities sources, with strong French-language depth (API overview). |
GET https://gallica.bnf.fr/SRU; auth: none. Rate: none published numerically in the API page. |
CQL explicitly distinguishes all (all words), any, adj, and Boolean and/or/not; emit gallica all "..." under the API docs. Date: dc.date comparisons on variable source precision or day-granular indexationdate=YYYYMMDD. Pagination: startRecord plus maximumRecords default 15/max 50; no total ceiling published. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes for BnF-produced descriptive metadata only. The BnF data portal places metadata under the French Open License for all reuse, including commercial reuse, with source attribution; digitized documents follow stricter Gallica terms. Filter provenance adj "bnf.fr", retain metadata only, record the retrieval/update date, and cite Bibliothèque nationale de France. Partner metadata, images, OCR, and document content stay outside the brief. |
DECLINED for Delegation 08 (2026-09-06). BnF-only SRU search returned HTTP 403. BnF metadata conditions also require source and retrieval date to be retained. The frozen normalized-result schema carries no retrieval timestamp. The source-specific corpus first_seen_at records insertion, potentially after other searches, and the report projection omits it. Static catalog credit alone cannot discharge the retrieval-date obligation. |
OSF API v2 registrations / osf — preregistrations/registered reports and connected research objects (official API spec). |
GET https://api.osf.io/v2/registrations/; auth: none for public resources or optional free token. Rate: 100/hour unauthenticated, 10,000/day authenticated in the spec. One survey probe of https://api.osf.io/v2/ returned 403; it was not retried. |
Resource filters use substring matching; the documented cross-object web Search syntax is not a published API contract. Registrations support modified-date comparisons, UTC precision undocumented. JSON:API page links, max 10 entities/response, no total ceiling in the API spec. | Q1 yes; Q2 yes; Q3 yes; Q4 yes for public, appropriately licensed objects. COS Terms §§5–6 define permitted use to include reproduce/store/transmit and grant a perpetual worldwide license subject to item/privacy restrictions. | REJECT. Compatible terms, but no stable documented cross-object keyword-search API and the single root probe returned 403. |
Figshare API v2 / figshare — datasets, theses, posters, presentations, code, and other research objects (article types). |
POST https://api.figshare.com/v2/articles/search; auth: none for public data. No automatic numeric limit, but Figshare recommends ≤1 request/second and may block abuse under Rate limiting. |
Bare spaces are OR with relevance ranking; explicit Boolean/quotes/grouping are supported in the search docs. Date: lower-bound-only published_since/modified_since, ISO date/time. Pagination: page/size or offset/limit, search window capped at offset 1,000 in the legacy feature reference. |
Although metadata policy applies CC0, the controlling Service Terms include APIs, limit use to personal noncommercial research, and prohibit using the service to power third-party search/products without authorization. Q1–3 no absent written permission; Q4 has no refresh duty but cannot cure that. | REJECT. Provider policy and service terms conflict; the product/search prohibition controls absent written authorization. |
CourtListener REST API v4.7 / courtlistener — U.S. case law, dockets, filings, and oral arguments (API overview). |
GET https://www.courtlistener.com/api/rest/v4/search/; auth: paid/contractual for product use because the current terms require product/team users to discuss a commercial agreement. Default authenticated rates: 5/minute, 50/hour, 125/day. |
Default connector AND; advanced Boolean/phrases/ranges in the query guide. Date: dateFiled:[YYYY-MM-DD TO YYYY-MM-DD], day. Pagination: opaque cursor; v4 removed the prior 100-page limit, but page size/total ceiling are undocumented in the migration guide. |
Q1 yes in principle; Q2 undocumented; Q3 no without commercial agreement; Q4 no refresh duty found. Credentials cannot be pooled/transferred and derived displays must not imply FLP endorsement under the terms. | REJECT. Paid/contractual product access is outside this survey’s adoption boundary. |
FAO AGRIS Open Data Set / agris — multilingual agriculture records spanning articles, datasets, reports, theses, books, and conference material, including underserved regions (ODS). |
Bulk archive at https://agris.fao.org/agris_ods/; auth: none. Rate: none published. No public keyword-search API is documented. |
Keyword combination, server-side date filtering/granularity, pagination, and query ceiling: undocumented/not applicable because the official download page exposes bulk ZIP/XML/RDF rather than a search endpoint. | Q1 yes; Q2 yes; Q3 yes; Q4 yes with attribution/provenance. FAO pages conflict between CC BY 4.0 on the current ODS page and CC BY 3.0 IGO on Download; both permit reuse with attribution, but the version mismatch must remain visible. | REJECT. Friendly terms, but no documented bounded keyword/date/pagination API suitable for resmon. |
PatentsView PatentSearch / patentsview — disambiguated U.S. granted-patent metadata/entities (reference). |
Legacy https://search.patentsview.org/api/v1/patent/ required a free X-Api-Key and allowed 45/minute, but new keys are suspended. USPTO paused PatentsView during ODP migration; successor auth requires a free USPTO.gov account, with rate unpublished. |
Explicit _text_all/_text_any/_text_phrase and _and/_or; date comparisons on patent_date=YYYY-MM-DD; cursor o.after plus o.size default 100/max 1,000, no total ceiling in the legacy reference. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes with CC BY attribution under the PatentSearch 2.3.2 release. | REJECT FOR NOW. The official API is paused and cannot onboard users; reassess only after ODP publishes its replacement contract. |
EPO Open Patent Services / epo_ops — worldwide patent bibliographic, legal-status, and text/image data (OPS overview). |
GET https://ops.epo.org/3.2/rest-services/published-data/search; auth: free registration/OAuth up to 4 GB/week, then paid. Fair-use limits are about 1 Mbit/s and 4 GB/week, not a fixed call count, under the Fair Use Charter. |
CQL requires explicit Boolean; bare spaces are undocumented, per the search-format FAQ. Publication dates support year/month/day and comparisons. Pagination: Range default 1–25/max 100; only 2,000 hits retrievable per query in the reference guide. |
OPS Terms §§3 and 7.1 permit own databases/products but require notified corrections to be incorporated without undue delay. Q1 yes; Q2 undocumented; Q3 conditional; Q4 no. | REJECT. Mandatory correction ingestion directly conflicts with no refresh after first retrieval. |
CiNii Research OpenSearch v2 / cinii — Japanese articles, books, dissertations, datasets, projects, and researchers (specification). |
GET https://cir.nii.ac.jp/opensearch/v2/all; auth: free application ID, with prior contact for commercial/other purposes under registration. Rate: none published; high-volume bursts may be blocked. |
Free-word q combines terms by AND. Date: from/until in YYYY or YYYYMM, month granularity. Pagination: count default 20/max 200 and start max 10,000, all in the OpenSearch spec. |
Academic Content Regulations Arts. 4–5 limit use to the user’s own academic research, restrict copying, and prohibit storage on a server shareable by others; API Regulations restrict key transfer/sublicensing. Q1 no safely; Q2 no; Q3 no; Q4 no refresh duty found. | REJECT. Cloud/storage and academic-purpose/key restrictions conflict with the product. |
Library of Congress JSON API / loc — U.S. books, periodicals, legislation, manuscripts, maps, and cultural-heritage collections (endpoint guide). |
GET https://www.loc.gov/search/?fo=json; auth: none. Rate: 20 requests/minute under Working Within Limits. |
q searches metadata/full text, but multi-term combination is undocumented in parameters. Date: dates=start/end, usually year precision. Pagination: sp/c, default 25, recommended max 1,000/page, deep-paging limit 100,000 under limits. |
The Library says collection items may remain copyrighted and users must make collection/item-specific rights assessments in its rights FAQ. No blanket API-metadata grant covering indefinite/cloud copies was found. Q1–3 unresolved; Q4 no refresh duty found. | REJECT. Public API access is not a reusable rights grant across heterogeneous collections. |
BASE OAI-PMH / HTTP interface / base — global repository/journal aggregator with books, theses, conference items, reports, patents, and datasets (service page). |
Search interface requires a use-case application/free key under the service page; OAI-PMH at http://oai.base-search.net/oai is IP-restricted to approved noncommercial projects. Numeric rate not published on those accessible pages. |
OAI dynamic sets use Solr field:value, but space combination is undocumented. Date: OAI from/until means BASE harvest time; publication date can be a dynamic date or normalized year filter, with source precision variable. Pagination: resumptionToken and no stated total ceiling in the OAI docs. |
BASE explicitly allows approved noncommercial clients to integrate subsets into local indexes, but its public pages do not grant the four uses or address private cloud backup; source abstracts retain third-party rights. Q1 conditional; Q2 unresolved; Q3 conditional; Q4 no duty published. | REJECT. Institutional/IP approval plus missing durable/cloud terms; seek a written license before reconsideration. |
Google Books Volumes API / google_books — global book metadata and previews (API guide). |
GET https://www.googleapis.com/books/v1/volumes; auth: free API key for public-data applications. Rate: none published in the guide; project quotas are assigned by Google. |
Multiple unquoted terms default to AND; phrases and exclusions supported. No publication-date filter; orderBy=newest only sorts. Pagination: startIndex and maxResults default 10/max 40, with no total ceiling documented in the guide. |
Google API Terms §5(e) and §8 prohibit building databases/permanent copies or caching beyond headers and require deletion of stored/cached content at termination absent owner permission. Q1 no; Q2 no; Q3 no for permanent records; Q4 no. | REJECT. Permanent local records directly violate the API terms. |
SciELO ArticleMeta / scielo — Latin American/Caribbean journals and multilingual regional literature (official API landing). |
Historical endpoint https://articlemeta.scielo.org/api/v1/article/; auth: undocumented; rate: none published. Known dead end: earlier plain requests returned 403; the single 2026-08-30 recheck returned 404, not a usable API response. |
The ArticleMeta tooling is harvesting-oriented, and the official landing publishes no supported keyword-search contract. Keyword combination, date filtering/granularity, pagination, and ceiling: undocumented for a search client. | No provider-wide API/metadata terms were found on the official landing that establish Q1–Q4 for normalized records containing abstracts; article licenses vary. | REJECT. Required known dead end: blocked/obsolete plain endpoint, no documented keyword API, and no usable blanket terms. |
ChemRxiv public API / chemrxiv — chemistry preprints (official service). |
Apparent public endpoint https://chemrxiv.org/engage/chemrxiv/public-api/v1/items; auth, rate: undocumented. The one plain survey request returned 403. |
Keyword combination, date filtering/granularity, pagination, and ceiling: undocumented in first-party public API material found during this survey. | ChemRxiv items display heterogeneous item licenses, and no blanket metadata/abstract reuse grant was found on the service. Q1–3 unresolved; Q4 no published refresh duty found. | REJECT. Required known dead end: plain programmatic access returned 403, with no public contract sufficient for a client brief. |
Directory of Open Access Books (DOAB) OAI-PMH / doab — global open-access book/chapter metadata (metadata page). |
https://directory.doabooks.org/oai/request; auth: none; rate: none published. A prior probe did not respond; the one 2026-08-30 Identify recheck returned 200. |
OAI-PMH has no keyword-search operation. Date from/until filters record datestamps and pagination uses resumptionToken under the OAI-PMH protocol; repository granularity/ceiling are not published by DOAB. |
Q1 yes; Q2 yes; Q3 yes; Q4 yes for metadata, which DOAB makes CC0 with no restrictions in its FAQ; book content is separate. | REJECT. Required known dead end updated: endpoint now responds, but OAI harvesting cannot execute resmon keyword monitoring. OAPEN’s REST search is the actionable books candidate. |
J-STAGE WebAPI / jstage — Japanese journals, conference proceedings, and technical reports (WebAPI overview). |
GET https://api.jstage.jst.go.jp/searchapi/do?service=3; auth: none for noncommercial use; commercial use requires prior approval under the terms. Numeric rate not published. |
The v2.0 manual says multiple terms in one parameter are space-separated AND; article title, author, keyword, abstract, and full-text fields are available. Date: pubyearfrom/pubyearto, year granularity. Pagination: start/count, max 1,000 per response and traversal beyond 1,000; no total ceiling documented. |
The controlling Japanese Terms Art. 3(5)–(6) prohibit machine-readable server/cloud storage or caching for 24 hours or more and require always displaying the latest J-STAGE information; Art. 9 requires “Powered by J-STAGE.” Q1 no; Q2 no; Q3 conditional; Q4 no. | REJECT. Required known dead end: the 24-hour storage cap and mandatory refresh are directly incompatible. |
African Journals Online (AJOL) / ajol — Africa-wide journal aggregator, filling the clearest geographic gap (AJOL search). |
No documented public search API endpoint found on the first-party AJOL service; auth/rate therefore undocumented. The public HTML search is not an API contract. | The web search defaults bare terms to OR and documents explicit AND/NOT/phrases, but API semantics do not exist. Server-side API date filtering, pagination, and ceiling: undocumented. |
AJOL-hosted journals publish heterogeneous copyright/license notices, and no AJOL-wide metadata license or API terms were found. Q1–4 unresolved for an aggregator client. | REJECT. High-value gap, but no documented API or blanket metadata-use grant; record it so future work starts with provider contact rather than scraping. |
WHO Global Index Medicus / global_index_medicus — regional health indexes including African and Eastern Mediterranean literature (GIM portal/guide). |
No documented public API endpoint found; auth/rate undocumented. The GIM portal is a human search service. | The search guide documents fielded terms, phrases, Boolean, and publication DA:YYYYMM (month granularity), but no API pagination or result ceiling. |
No GIM metadata/API reuse terms establishing indefinite local/cloud retention were found on the portal. Q1–4 unresolved. | REJECT. Important Africa/Middle-East discovery target, but no public API contract or durable-copy grant. |
Recommended first implementation tranche
- Dryad — the cleanest dataset source: no key and CC0 throughout.
- GovInfo — the strongest law/official-report expansion, with a generous free key and exact search/date contracts.
- NDL Search — the strongest geographic/language correction; use only the open provider slice, show required credits, and partition around its 500-result ceiling.
- Open Library — a no-key books/humanities source whose CC0/public-domain provenance and caching guidance fit resmon’s metadata-only design.
NIST RMM is the next no-key brief. OAPEN should follow once its REST pagination behavior is established in a one-request implementation spike. DPLA, Europeana, and OpenReview are lower-priority free-key additions.