Research Impact Ledger
Research Impact Ledger
A self-hosted research-impact tracker for everything I build: software, benchmarks, datasets and technical writeups. It answers “when does someone use or cite my work?” without depending on any single service.
GitHub ──► CITATION.cff ──► Zenodo ──► DOI ──┐
├─► OpenAlex ─┐
└─► Crossref ─┤
├─► cache ─► /research-impact/ page
GitHub (stars/forks/clones/views) ──────────────────────────┤
Hugging Face (downloads/likes) ─────────────────────────────┤
Web mentions (curated / Brave) ─────────────────────────────┘
└─► email / Telegram / Discord alerts
Files
| Path | Role |
|---|---|
_data/research_impact.yml |
Source of truth. The ledger: artifacts, repos, DOIs, posts, licenses, detector settings. Edit this by hand. |
_data/research_impact_mentions.yml |
Manual web mentions, keyed by artifact id. |
_data/research_impact_cache.json |
Generated metrics (committed by CI). |
_data/research_impact_state.json |
What we have already alerted on (committed by CI). |
scripts/research_impact/* |
Detectors, collector and notifier. |
_pages/research-impact.md |
The public page at /research-impact/. |
Running locally
pip install -r scripts/research_impact/requirements.txt
GH_TOKEN=$(gh auth token) python3 scripts/research_impact/collect.py --no-notify
Then rebuild the site (bundle exec jekyll serve) to see /research-impact/.
Adding a new artifact
- Add an entry under
artifacts:in_data/research_impact.yml(copy an existing block).idmust be unique and URL-safe. - Run the collector to pick up usage signals.
- Mint a DOI (see below) and the ledger is updated for you.
Identity (ORCID)
owner_orcid in _data/research_impact.yml is the one identifier that ties
everything together. It is stamped into each repo’s CITATION.cff and
.zenodo.json, so every deposit is attributed to the same person. That is what
lets ORCID import the works (via ORCID -> Add works -> Search & link ->
Zenodo, or by authorizing Zenodo as a trusted organization for automatic
updates) and what OpenAlex uses to disambiguate the author.
When adding a new artifact, put the ORCID in its .zenodo.json creators:
"creators": [{ "name": "Singh, Yuvraj", "orcid": "0009-0005-4120-555X" }]
Minting a DOI automatically
scripts/research_impact/mint_doi.py uses the Zenodo REST API directly, so it
works without enabling the Zenodo<->GitHub webhook. It reads the repo’s
committed .zenodo.json, creates a deposition, uploads the GitHub source
archive for the chosen ref, publishes it and writes the DOI back into the
ledger.
Create a token at
https://zenodo.org/account/settings/applications/tokens/new/ with both
deposit:write and deposit:actions scopes, then:
export ZENODO_TOKEN=...
# Dry run: creates a draft, reports it, then deletes it. Mints nothing.
python3 scripts/research_impact/mint_doi.py --artifact smoltorrent
# Full rehearsal against the sandbox (throwaway DOI, safe to publish):
ZENODO_SANDBOX_TOKEN=... python3 scripts/research_impact/mint_doi.py \
--artifact smoltorrent --sandbox --publish
# Real DOI for a tagged release:
python3 scripts/research_impact/mint_doi.py --artifact smoltorrent --ref v1.0.0 --publish
# Every artifact that still has no DOI:
python3 scripts/research_impact/mint_doi.py --all --publish
Publishing is irreversible, which is why the default is a dry run. Prefer
archiving a git tag (--ref v1.0.0) over HEAD so the DOI points at a
frozen version.
Versioned DOIs
Tag a release, then mint a new version of the existing record so both DOIs are grouped under one concept DOI (which always resolves to the latest version):
gh release create v1.0.0 --target master --title "v1.0.0"
python3 scripts/research_impact/mint_doi.py \
--artifact smoltorrent --ref v1.0.0 --new-version v1.0.0 --publish
The ledger keeps doi: as the concept DOI, records version and
version_doi, and points zenodo_id at the newest record so the next
--new-version run works.
You can also run this from GitHub: Actions -> Mint Zenodo DOI -> Run
workflow (add ZENODO_TOKEN/ZENODO_SANDBOX_TOKEN as repository secrets
first). The workflow refreshes the cache and commits the DOI for you.
Detectors
- OpenAlex — for each DOI, lists every citing work (title, authors, venue,
year, DOI) via the
cites:filter. - Crossref —
is-referenced-by-countper DOI. - GitHub — stars, forks, watchers, contributors, latest release, and (with a token that can read them) 14-day views/clones.
- Hugging Face — downloads and likes for datasets, models and collections.
- Web mentions — curated entries, or the Brave Search API when
settings.mentions: braveandBRAVE_API_KEYis set.
Google Scholar has no public API, so it is not queried directly. The Zenodo DOI is what makes Scholar, OpenAlex and Crossref able to attribute a citation to the software or dataset.
Alerts
Every collector run diffs against _data/research_impact_state.json and alerts
on new citations (and gentle star/download milestones). Configure any of:
| Channel | Environment variables |
|---|---|
| Telegram | RESEARCH_IMPACT_TELEGRAM_BOT_TOKEN, RESEARCH_IMPACT_TELEGRAM_CHAT_ID |
| Discord | RESEARCH_IMPACT_DISCORD_WEBHOOK |
| Email (Resend) | RESEARCH_IMPACT_RESEND_KEY, RESEARCH_IMPACT_EMAIL_FROM, RESEARCH_IMPACT_EMAIL_TO |
| Email (SMTP) | RESEARCH_IMPACT_SMTP_HOST, _PORT, _USER, _PASS, RESEARCH_IMPACT_EMAIL_FROM, RESEARCH_IMPACT_EMAIL_TO |
| Generic webhook | RESEARCH_IMPACT_WEBHOOK_URL |
Add these as repository secrets; .github/workflows/research-impact.yml runs
daily and commits the refreshed cache.