158 science skills turn Claude Code into a researcher

Scientific agent skills pack 158 research workflows and wire up more than 100 scientific databases for any coding agent. The database layer does the heavy lifting here, since a general agent already writes Python against any library but has no way to know which database holds the answer or how to query it correctly.

Key Takeaways

  • 158 packaged workflows cover biology, chemistry, medicine, and drug discovery.
  • Over 100 scientific databases arrive pre-wired, which saves the worst of the work.
  • Works with Claude Code, Cursor, Codex, and other agents on the open standard.
  • One benchmark found the skill cut tool errors by 92 percent.
  • Small models mostly ignore the skill unless you name it in the prompt.

What scientific agent skills add that a general agent lacks

A capable coding agent can import a scientific Python library and call a REST API today. Give it the docs and it will write something that runs.

Choosing is the unreliable part. Six overlapping tools exist for most bioinformatics jobs, and picking the wrong one wastes a day. The agent also has no feel for which parameters change the answer in a field it has never worked in.

A skill supplies that missing judgment. Each one is a folder with a SKILL.md file holding curated documentation, worked examples, and the conventions of a field. The agent loads it only when the task calls for it.

The Scientific Agent Skills project frames it the same way. Its own README admits the agent can use any Python package by itself. Named skills just make it stronger and steadier on these jobs.

Skills also encode which source to query and the syntax it expects, something no model picked up well on its own. Without that, an agent will invent a pipeline that looks fine and breaks at step one.

The measured payoff, and where it vanishes

K-Dense ran a controlled study on one of its own skills, and the numbers are more useful than its marketing. The pyOpenMS skill benchmark covers 250 runs across ten mass-spectrometry tasks: five repeats per arm, two models, an automated grader, and no LLM judge.

On a strong model the skill helped on every axis at once.

Measure (Claude Sonnet)Without the skillWith the skill
Task success96% (48/50)100% (50/50)
pyOpenMS API errors per run1.000.08
Wall-clock time per task42.3 s33.8 s
Cost per correct result$0.180$0.162
Runs that invoked the skilln/a98%

The API error rate fell by 92 percent, and output tokens dropped by more than a third. The mechanism is dull and believable: a strong model that hits a changed API can usually recover, but it burns turns and tokens doing it. The skill gets the call right the first time and skips the debugging loop.

On version-specific pipeline tasks the skill was about twice as fast and 24 percent cheaper. On simple jobs like a tryptic digest it was mild overhead, since a strong model was already fine.

Claude Haiku without the skill took two minutes per task, made nearly four API errors per run, and timed out on the hardest jobs. Installing the skill barely helped, because the model reached for it in only 8 percent of runs.

Measure (Claude Haiku)No skillSkill availableSkill forced
Success rate74%86%88%
Runs that invoked the skill0%8%100%
Wall-clock time per task121.5 s83.5 s31.1 s
Cost per task$0.152$0.133$0.074
API errors per run3.743.000.52

Forcing the skill with a one-line instruction made the cheap model the cheapest arm in the whole study, at $0.074 per task.

Scatter chart plotting task success rate against cost per task, with forced-skill Haiku at 88 percent success for 7.4 cents while Sonnet reaches 100 percent at 16 cents
Forced to use the skill, the cheap model lands at the lowest cost of any arm in the study
Image: K-Dense

a great skill is necessary but not sufficient, because a weak model has to actually trigger it

K-Dense

Two caveats belong with those numbers. The Haiku cost figures in the first two columns exclude runs that hit a ten-minute timeout, so the real gap is wider. And the study tests one skill, so the result doesn’t automatically transfer to the other 157.

The failures that should worry a working scientist are the ones that return a clean number that happens to be wrong.

K-Dense

The study caught two of them in the no-skill runs. One computed a peptide mass correctly but reported the wrong charged m/z, 574.26 instead of 583.27, a number that would sail downstream unnoticed. Another ran a full adduct-grouping pipeline without a single error and annotated zero adducts, because it used the wrong syntax. On feature detection the weak model produced noise-inflated results in 13 of 15 runs without the skill, and 0 of 15 with it forced.

Paired charts: a bar chart of pyOpenMS API errors per run falling from 3.74 to 0.52 on Haiku, and a log-scale dot plot where no-skill runs report tens of thousands of features against a clean target near 390
Left, the API errors. Right, the runs that finish cleanly and report thousands of features that are mostly noise
Image: K-Dense

Still, the effect has a ceiling. A separate K-Dense study on NVIDIA BioNeMo skills reached a narrower verdict. Skills help most with routing to odd endpoints and with weak models. They don’t make the science model itself any more accurate.

What the 158 skills actually cover

Counting the directories in the repo confirms the badge: 158 skill folders, each with its own SKILL.md, grouped by field.

DomainExample workflows
Bioinformatics and genomics (26 skills)Sequence analysis, single-cell RNA-seq, gene regulatory networks, variant annotation, RNA velocity
Scientific communication (27 skills)Literature review, evidence-traceable writing, peer review, posters, slides, citation management
Data analysis and visualization (22 skills)Statistics, network analysis, time series, publication-quality figures
Machine learning and AI (14 skills)Deep learning, reinforcement learning, model interpretability, Bayesian methods
Research methodology (13 skills)Hypothesis generation, grant writing, scenario analysis, scholar evaluation
Infrastructure and platforms (11 skills)Modal cloud compute, GPU acceleration, Nextflow, DNAnexus, LatchBio, OMERO
Cheminformatics and drug discovery (10 skills)Property prediction, virtual screening, ADMET, docking, lead optimization
Databases and data access (10 skills)100+ databases, DepMap, PrimeKG, live pathogen variant surveillance
Clinical research (8 skills)Trials, pharmacogenomics, PK/PD modeling, dose selection
Proteomics and mass spectrometry (2 skills)LC-MS/MS processing, peptide identification, spectral matching

Smaller clusters cover materials science, physics and astronomy, medical imaging, neuroscience, protein engineering, lab automation, geospatial science, and regulatory standards.

Find the two or three skills matching workflows you already run by hand, open those SKILL.md files, and read them. If they encode things you know to be true, the rest of the library probably earns its place.

The database layer is the real payload

Anyone can write a skill file. Wiring up more than 100 scientific databases with the right query rules is the part you can’t copy in an afternoon.

The bulk of it sits in one database-lookup skill. That single file covers 78 public databases across chemistry, genomics, clinical data, patents, and economics. PubChem, ChEMBL, UniProt, PDB, ClinVar, COSMIC, ClinicalTrials.gov, FDA, FRED, and USPTO are all in there. Dedicated skills add DepMap, PrimeKG, and pathogen surveillance. Packages like BioServices and BioPython push the total past 100.

Diagram contrasting a grid of 78 separate always-loaded skill files on the left with a single highlighted router skill on the right that branches out to greyed-out reference files loaded on demand
Seventy-eight standing skill files on the left, one router with on-demand references on the right
Image: K-Dense

Every database has its own identifiers, query syntax, rate limits, and idea of what a record is. A guessed query returns nothing, and the agent may report that as a real negative result, while a skill that encodes the right pattern at least fails loudly.

Bundling them into one skill has a measured payoff. K-Dense tested one skill against 78 separate ones. The split version costs 13.9 times more always-on context: 3,358 tokens per request against 242. The reference corpus runs to about 100,000 tokens, and 93 percent of it never loads until the agent picks a database.

Two bar charts: always-on system prompt tokens at 242 for one consolidated skill against 3,358 for 78 separate skills, and a stacked bar showing 93.1 percent of the 100,298-token reference corpus lazy-loaded
The standing context tax on the left, the corpus that stays out of context on the right
Image: K-Dense

Routing quality held up too, with a wrinkle. On single-database lookups all five tested models scored 96 to 100 percent either way, so the guide bought nothing. Cross-domain questions told a different story: without the guide one model dropped to 89 percent and another to 63 percent, while with it all five held at 100 percent.

One practical detail the docs bury: most of these databases are open. I read all 78 reference files, and only 19 need an API key. Those skew toward economics, government, and weather sources such as FRED, BEA, Eurostat, and Materials Project. The core biology and chemistry stack is keyless: PubChem, UniProt, ChEMBL, PDB, and Ensembl. You can run most of the library without signing up anywhere.

The same team also ships K-Dense BYOK , a free desktop app that bundles these skills with a research workspace and your own model keys.

How to add scientific agent skills to your coding agent

Install the library, confirm your agent can see it, then run a workflow and check what it actually did.

Pick your agent

The library targets the open Agent Skills standard. That means Claude Code , Cursor , Codex, Gemini CLI, and Google Antigravity all work from the same source.

Install only the domains you use

The standard installer is npx skills add K-Dense-AI/scientific-agent-skills. If you use the GitHub CLI , gh skill install K-Dense-AI/scientific-agent-skills picks the right directory for your host and records provenance metadata. Pin a release with --pin v2.62.0 rather than tracking a branch. Skip the full collection. The project’s own security notice advises against it, since many skills now come from the community, and 158 skills add up to a lot of standing context.

Verify discovery

Start a new session and ask the agent to list the skills it can see. If the domain skills are missing, the install path is wrong for your build. In Cursor, check Settings then Rules.

Start with a read-only task

Ask for a database lookup rather than an analysis. A single fact from PubChem or UniProt is something you can check against the source in seconds.

Name the skill in your prompt

If you run a small or cheap model, say “use the pyopenms skill” instead of hoping it triggers. Auto-discovery worked 98 percent of the time on a strong model and 8 percent on a weak one.

Read what it ran

Ask the agent to show the commands and the sources it used. A skill is documentation plus examples, so the underlying call should be inspectable. Treat any result you can’t inspect as unverified.

Check the result against the primary source

Open the database record it cited. In scientific work this step isn’t optional. A pipeline that annotates zero adducts and exits cleanly will never tell you it failed.

Keep the environment separate

Scientific Python stacks clash often. The repo wants Python 3.13 and uv for its own tooling. Give the agent its own environment, so a dependency change doesn’t break your lab’s setup.

What the rename says about agent skills

The project used to be called Claude Scientific Skills and is now Scientific Agent Skills, with the same content pointed at the open Agent Skills standard.

That change goes past branding. Write a skill once and Claude Code, Cursor, Codex, and Antigravity can all load it. That is a different business case from a per-vendor plugin. The Model Context Protocol set the standard for how agents reach tools. The skills standard is trying to do the same for packaged know-how.

The engineering around the library is unusually serious for a skill collection. It’s MIT licensed and sits on a 2.x release line. Five GitHub Actions workflows run on it, including a pull-request skill scan, a spec validator, and a test suite. Every skill shipping a scripts/ folder needs tests, and CI blocks pull requests that skip them.

A skill is instructions your agent will follow. Installing one is a trust decision about its authors. K-Dense scans the repo weekly with the Cisco AI Defense skill scanner and publishes what it finds. Its own guide to reviewing skills before you install says to read the full SKILL.md and any bundled scripts first. You can run that scanner yourself too.

Individual skills carry their own licences, which can differ from the repo’s MIT terms. Check the license field in each SKILL.md if you plan commercial use.

The promise has a hard limit. Skills make an agent better at scientific workflows, and the output still needs a domain expert reading it with a normal amount of doubt.