← cd ..

$ cron.daily paper_ranker.py # 2026-09-19

// live, automated · updated 19 Sept 2026

Every morning a GitHub Actions cron pulls the newest papers from cs.AI, cs.CL, cs.LG, cs.MA on arXiv, scores each one with Jev (TypeSafe AI's typed-decision model — no text generation, just calibrated yes/no + score answers), and ranks them.

// ranked against:

"LLM agents, tool use, and structured/typed model outputs"

// priority = 0.7 × worth-reading confidence + 0.3 × novelty score — both from Jev, weighting decided in code

$ sort -k1 -rn ./papers/2026-09-19.json

01. Quantifying Overclaiming Propensity in Frontier LLM Agents

87% worth reading · novelty 1.8/3 · benchmark/eval · cs.AI

02. FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

86% worth reading · novelty 1.8/3 · new method · cs.MA

03. Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

85% worth reading · novelty 1.8/3 · new method · cs.CL

04. Reputation as Community Memory for the Agentic Web

83% worth reading · novelty 1.9/3 · new method · cs.MA

05. Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights

83% worth reading · novelty 1.8/3 · new method · cs.AI

06. Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

81% worth reading · novelty 1.6/3 · new method · cs.AI

07. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

82% worth reading · novelty 1.6/3 · new method · cs.AI

08. SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

76% worth reading · novelty 1.9/3 · benchmark/eval · cs.MA

09. Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

79% worth reading · novelty 1.7/3 · new method · cs.CL

10. An Empirical Study of Harness Design for Coding Agents

86% worth reading · novelty 1.2/3 · benchmark/eval · cs.AI

11. RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

78% worth reading · novelty 1.7/3 · new method · cs.AI

12. Large Language Models as Falsifiers for Cyber-Physical Systems

74% worth reading · novelty 1.9/3 · new method · cs.AI

13. Language-model groups overstate consensus when replaying human deliberation on a reasoning task

74% worth reading · novelty 1.7/3 · benchmark/eval · cs.MA

14. Agentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks

70% worth reading · novelty 1.9/3 · new method · cs.MA

15. LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents

76% worth reading · novelty 1.4/3 · new method · cs.MA

16. Flag Game: A Toy Model for Mechanistic Swarm Interpretability

68% worth reading · novelty 1.9/3 · new method · cs.MA

17. On-Demand Attention: Language Models Know When to Recall

70% worth reading · novelty 1.8/3 · new method · cs.CL

18. Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions

81% worth reading · novelty 1.0/3 · benchmark/eval · cs.MA

19. Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

80% worth reading · novelty 1.0/3 · new method · cs.MA

20. Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

68% worth reading · novelty 1.7/3 · new method · cs.AI