Research Impact Ledger

Research Impact Ledger

A self-hosted research-impact tracker for everything I build: software, benchmarks, datasets and technical writeups. It answers “when does someone use or cite my work?” without depending on any single service.

GitHub ──► CITATION.cff ──► Zenodo ──► DOI ──┐
                                              ├─► OpenAlex ─┐
                                              └─► Crossref ─┤
                                                            ├─► cache ─► /research-impact/ page
GitHub (stars/forks/clones/views) ──────────────────────────┤
Hugging Face (downloads/likes) ─────────────────────────────┤
Web mentions (curated / Brave) ─────────────────────────────┘
                                                            └─► email / Telegram / Discord alerts

Files

Path Role
_data/research_impact.yml Source of truth. The ledger: artifacts, repos, DOIs, posts, licenses, detector settings. Edit this by hand.
_data/research_impact_mentions.yml Manual web mentions, keyed by artifact id.
_data/research_impact_cache.json Generated metrics (committed by CI).
_data/research_impact_state.json What we have already alerted on (committed by CI).
scripts/research_impact/* Detectors, collector and notifier.
_pages/research-impact.md The public page at /research-impact/.

Running locally

pip install -r scripts/research_impact/requirements.txt
GH_TOKEN=$(gh auth token) python3 scripts/research_impact/collect.py --no-notify

Then rebuild the site (bundle exec jekyll serve) to see /research-impact/.

Adding a new artifact

  1. Add an entry under artifacts: in _data/research_impact.yml (copy an existing block). id must be unique and URL-safe.
  2. Run the collector to pick up usage signals.
  3. Mint a DOI (see below) and the ledger is updated for you.

Identity (ORCID)

owner_orcid in _data/research_impact.yml is the one identifier that ties everything together. It is stamped into each repo’s CITATION.cff and .zenodo.json, so every deposit is attributed to the same person. That is what lets ORCID import the works (via ORCID -> Add works -> Search & link -> Zenodo, or by authorizing Zenodo as a trusted organization for automatic updates) and what OpenAlex uses to disambiguate the author.

When adding a new artifact, put the ORCID in its .zenodo.json creators:

"creators": [{ "name": "Singh, Yuvraj", "orcid": "0009-0005-4120-555X" }]

Minting a DOI automatically

scripts/research_impact/mint_doi.py uses the Zenodo REST API directly, so it works without enabling the Zenodo<->GitHub webhook. It reads the repo’s committed .zenodo.json, creates a deposition, uploads the GitHub source archive for the chosen ref, publishes it and writes the DOI back into the ledger.

Create a token at https://zenodo.org/account/settings/applications/tokens/new/ with both deposit:write and deposit:actions scopes, then:

export ZENODO_TOKEN=...

# Dry run: creates a draft, reports it, then deletes it. Mints nothing.
python3 scripts/research_impact/mint_doi.py --artifact smoltorrent

# Full rehearsal against the sandbox (throwaway DOI, safe to publish):
ZENODO_SANDBOX_TOKEN=... python3 scripts/research_impact/mint_doi.py \
    --artifact smoltorrent --sandbox --publish

# Real DOI for a tagged release:
python3 scripts/research_impact/mint_doi.py --artifact smoltorrent --ref v1.0.0 --publish

# Every artifact that still has no DOI:
python3 scripts/research_impact/mint_doi.py --all --publish

Publishing is irreversible, which is why the default is a dry run. Prefer archiving a git tag (--ref v1.0.0) over HEAD so the DOI points at a frozen version.

Versioned DOIs

Tag a release, then mint a new version of the existing record so both DOIs are grouped under one concept DOI (which always resolves to the latest version):

gh release create v1.0.0 --target master --title "v1.0.0"
python3 scripts/research_impact/mint_doi.py \
    --artifact smoltorrent --ref v1.0.0 --new-version v1.0.0 --publish

The ledger keeps doi: as the concept DOI, records version and version_doi, and points zenodo_id at the newest record so the next --new-version run works.

You can also run this from GitHub: Actions -> Mint Zenodo DOI -> Run workflow (add ZENODO_TOKEN/ZENODO_SANDBOX_TOKEN as repository secrets first). The workflow refreshes the cache and commits the DOI for you.

Detectors

  • OpenAlex — for each DOI, lists every citing work (title, authors, venue, year, DOI) via the cites: filter.
  • Crossref — is-referenced-by-count per DOI.
  • GitHub — stars, forks, watchers, contributors, latest release, and (with a token that can read them) 14-day views/clones.
  • Hugging Face — downloads and likes for datasets, models and collections.
  • Web mentions — curated entries, or the Brave Search API when settings.mentions: brave and BRAVE_API_KEY is set.

Google Scholar has no public API, so it is not queried directly. The Zenodo DOI is what makes Scholar, OpenAlex and Crossref able to attribute a citation to the software or dataset.

Alerts

Every collector run diffs against _data/research_impact_state.json and alerts on new citations (and gentle star/download milestones). Configure any of:

Channel Environment variables
Telegram RESEARCH_IMPACT_TELEGRAM_BOT_TOKEN, RESEARCH_IMPACT_TELEGRAM_CHAT_ID
Discord RESEARCH_IMPACT_DISCORD_WEBHOOK
Email (Resend) RESEARCH_IMPACT_RESEND_KEY, RESEARCH_IMPACT_EMAIL_FROM, RESEARCH_IMPACT_EMAIL_TO
Email (SMTP) RESEARCH_IMPACT_SMTP_HOST, _PORT, _USER, _PASS, RESEARCH_IMPACT_EMAIL_FROM, RESEARCH_IMPACT_EMAIL_TO
Generic webhook RESEARCH_IMPACT_WEBHOOK_URL

Add these as repository secrets; .github/workflows/research-impact.yml runs daily and commits the refreshed cache.

Telegram setup (3 steps) 1. In Telegram, message [@BotFather](https://t.me/BotFather) -> `/newbot`, pick a name, and copy the bot token (`123456:ABC-...`). 2. Send any message to your new bot, then open `https://api.telegram.org/bot/getUpdates` and copy `result[0].message.chat.id` (a number, may be negative for groups). 3. Add both as repository secrets: ```bash gh secret set RESEARCH_IMPACT_TELEGRAM_BOT_TOKEN --repo YuvrajSingh-mist/SmolHub-Website gh secret set RESEARCH_IMPACT_TELEGRAM_CHAT_ID --repo YuvrajSingh-mist/SmolHub-Website ``` Test it without waiting for a real citation: ```bash RESEARCH_IMPACT_TELEGRAM_BOT_TOKEN=... RESEARCH_IMPACT_TELEGRAM_CHAT_ID=... \ python3 -c "from scripts.research_impact import notify; \ notify.send_telegram('Research Impact test alert')" ``` </details> > **Traffic metrics:** the default `GITHUB_TOKEN` can only read traffic for the > repository that owns the workflow. To track views/clones for the *other* > repos, create a fine-grained PAT and add it as the `RESEARCH_IMPACT_TOKEN` > secret: > > 1. Go to <https://github.com/settings/personal-access-tokens/new>. > 2. **Token name:** `research-impact-traffic`. > 3. **Expiration:** e.g. 1 year (calendar a reminder to rotate). > 4. **Repository access:** "Only select repositories" -> `smoltorrent`, > `smolcluster`, `smolperfbenchmark`, `AndroidLife` (and any others you add). > 5. **Permissions -> Repository permissions -> Administration: Read-only.** > That is the permission that unlocks the traffic endpoints. Nothing else is > needed (leave Contents/Metadata at their defaults). > 6. Generate, copy the `github_pat_...` value. > 7. In the website repo: **Settings -> Secrets and variables -> Actions -> New > repository secret**, name `RESEARCH_IMPACT_TOKEN`, paste the token. > > The collector then reads > `GET /repos/{owner}/{repo}/traffic/views` and `.../traffic/clones` for each > artifact. If the secret is absent, those two fields are simply omitted and > everything else still works. > > To rotate, generate a new token and update the secret; the old one can be > revoked at <https://github.com/settings/tokens?type=beta>.