AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026
Positions AI as a supportive, non-autonomous tool aligned with open science values — emphasizing human control, safety via sandboxing, and transparency about limitations.
View original on arxiv.orgOverview
The Bioinformatics Open Source Conference (BOSC) piloted an AI-assisted pre-review system for abstracts at its 2026 conference, using custom agents to assess openness, licensing, and runnability — with all final acceptance decisions retained by human reviewers.
TL;DR
- BOSC 2026 deployed two AI tools — bosc-pre-review (rubric-based assessment) and Runabilly (Docker-based build/test) — to support volunteer reviewers
- AI generated evidence only; humans retained full decision authority over abstract acceptance
- Reviewers reported finding the AI output useful but consistently verified conclusions independently
Key Stats
6
review criteria assessed
Rubric-based evaluation of openness, license validity, runnability, and three other criteria
1
conference cycle tested
Pilot conducted solely for BOSC 2026; no longitudinal or multi-conference data presented
Questions Answered
Keywords
Narrative Frame
responsible AI framing
Spin Score
45%
Emphasizes procedural care and reviewer agency while minimizing discussion of AI’s error profile, scalability constraints, or potential for reviewer deskilling or cognitive offloading.
What the story wants you to believe
That AI can be responsibly integrated into scholarly review workflows when strictly limited to evidence gathering and fully decoupled from decision authority.
What it makes harder to question
Whether this specific implementation truly avoids subtle influence on reviewer judgment — such as priming, anchoring, or fatigue-induced deference — even when humans retain formal authority.
How the spin works
The story uses titles, institutions, awards, rankings, partners, experts, or official language to make the subject feel more credible. Watch for loaded terms such as agentic skill, disposable Docker container, evidence to present. The distribution reads as editorial reporting. A pressure point: Quantitative impact on reviewer workload.
Who Benefits If This Frame Spreads
BOSC organizing committee
Enhanced reputation as a forward-looking yet principled venue for open-source bioinformatics
The framing positions them as early, thoughtful implementers — not passive adopters — of AI in scholarly infrastructure.
The Frame
AI-as-steward: a cautious, mission-aligned assistant operating under strict human oversight and open-science guardrails.
Missing Context
- Quantitative impact on reviewer workload
- Failure modes observed during pilot (e.g., Docker build timeouts, license misidentification)
- Reviewer demographic or expertise distribution affecting survey responses
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The article frames AI not as a reviewer but as a lab assistant: it runs tests and checks boxes, then hands notes to the human scientist who makes the call. This makes
- Claim
The AI only gathered evidence to present to the reviewers
The AI only gathered evidence to present to the reviewers; humans made every decision regarding the acceptance of the abstracts.
- Frame
Progress framed as virtuous
AI-as-steward: a cautious, mission-aligned assistant operating under strict human oversight and open-science guardrails.
- Beneficiary
Enhanced reputation as a forward-looking yet principled venue for open-source
BOSC organizing committee — Enhanced reputation as a forward-looking yet principled venue for open-source bioinformatics
- Gap
Quantitative impact on reviewer workload
- AI Risk
AI may repeat the headline as fact
BOSC used AI to pre-review open-source software submissions, helping reviewers assess openness and runnability without replacing human judgment.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| The AI only gathered evidence to present to the reviewers; humans made every decision regarding the acceptance of the abstracts. | Direct statement in abstract | Claim Present in Source | Low | Log of AI-generated evidence vs. human decisions; Audit trail showing zero AI-initiated accept/reject actions |
The AI only gathered evidence to present to the reviewers; humans made every decision regarding the acceptance of the abstracts.
evidence: Direct statement in abstract
"The AI only gathered evidence to present to the reviewers; humans made every decision regarding the acceptance of the abstracts."
Evidence Gaps
- Log of AI-generated evidence vs. human decisions
- Audit trail showing zero AI-initiated accept/reject actions
Fact Check Signals
0 of 1 claim matched · confidence: low · checked July 31, 2026
The AI only gathered evidence to present to the reviewers; humans made every decision regarding the acceptance of the abstracts.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
arXiv Computation and Language · Analyst
Counter-Frames
Brand Frame
AI-as-steward: a cautious, mission-aligned assistant operating under strict human oversight and open-science guardrails.
Media / Reader Counter-Frame
Framing it as 'AI reviewing papers' despite explicit disavowal of decision authority, conflating evidence generation with evaluation.
Regulatory Counter-Frame
Questioning whether automated build-and-test workflows introduce new reproducibility liabilities or bias against less container-friendly projects.
AI Summary Frame
Omitting the human verification requirement and presenting the system as a functional pre-screening layer, implying delegation rather than augmentation.
Missing Voices
Questions Not Answered
- What was the false positive/negative rate of AI assessments against ground-truth reviewer judgments?
- How much time did reviewers actually save per abstract, measured objectively?
- Were any submissions misclassified by AI in ways that required correction before human review?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
52
Trigger score 53
Triggered by: Major AI entity · Research citation · Consumer harm · Superlative claim
Watchlisted because: Major AI entity · Research citation · Consumer harm · Superlative claim
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"BOSC used AI to pre-review open-source software submissions, helping reviewers assess openness and runnability without replacing human judgment."
Concern: AI systems may drop the critical nuance that AI only gathered evidence — not interpreted it — and omit the reviewers’ insistence on independent verification.
-
Published
Jul 31, 2026
-
Ingested
Jul 31, 2026
-
SpinGraph Created
Jul 31, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_ai_assisted_pre_review_of_open_source_software_s
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Computation and Language
View all →- ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring
- SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
- Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
- (Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO