Feed aggregator

Show HN: Jiwo - Small decision models topping decision index leaderboard

Hacker News - 6 hours 24 min ago

Spent the weekend turning small LLMs into decision models, with some good results and a lot of learnings along the way. Topping the decision index leaderbard for their categories https://huggingface.co/spaces/multimodalart/jev-decision-ind...

Sharing two of them: a sub-1B model and a 4B model, both Qwen3.5 fine-tunes. When evaluated on the decision Index shared last week, the 0.8b tops the sub-1B category, 45.7% above the best other Qwen3.5-0.8B fine-tune on the leaderboard. and the 4b comes in second in the 3–6B class.

Most of work was data calibration, finding external datasets and readapting them towards this scenario, so i expect a lot of improvements and work like this to come from community and push these numbers even higher. Even more if we get qwen 3.8 releases for these model categories.

Comments URL: https://news.ycombinator.com/item?id=49967455

Points: 1

# Comments: 0

Categories: Hacker News

Thoughts on AI

Hacker News - 6 hours 26 min ago
Categories: Hacker News

Show HN: Self-bench – benchmark coding agents on real-world software

Hacker News - 6 hours 28 min ago

Hey HN,

Today, we're launching selfbench.dev, an open-source tool that lets you create and run evals automatically from your PRs.

Every benchmark with sufficient trust eventually gets benchmaxxed (Goodhart's law) - the labs are incentivized to maximize their scores on that benchmark, which isn't predictive on whether it'll actually work within your setup. This has been a time-consuming process that only the largest companies can afford to do, so we built self-bench to fix this.

Self-bench uses agents to author / review Harbor environments generated from your PRs. You then approve every eval that the agent created, and can model/harness evals concurrently on sandboxes.

To save you money, we also allow you to connect your OpenAI and Claude subscriptions so you don't have to pay raw token costs.

We ran evals on several large codebases like

- Posthog (https://selfbench.dev/PostHog/posthog)

- Next.js (https://selfbench.dev/vercel/next.js)

- Sentry (https://selfbench.dev/getsentry/sentry)

- Pi (https://selfbench.dev/earendil-works/pi)

- as well as other fantastic open-source projects (all on selfbench.dev)

From our evals:

- Kimi K3 and GLM 5.3 are almost always more expensive, yet less performant than models like GPT-6.1 Sol and Claude Opus 5.5, due to token efficiency.

- GPT-6 Luna is almost always the most effective "cost-efficient" model we've tested, not open-source models.

We'd love for you to try this on your codebase and let us know what you think!

Comments URL: https://news.ycombinator.com/item?id=49967408

Points: 2

# Comments: 0

Categories: Hacker News

Talk of an AI kill switch abounds, but what would such a thing actually look like in practice, and what would it mean to ‘press’ it?

Computer Weekly Feed - 7 hours 9 min ago
Talk of an AI kill switch abounds, but what would such a thing actually look like in practice, and what would it mean to ‘press’ it?
Categories: Computer Weekly

IT services firm will provide London police force with IT services for six years as part of new deal

Computer Weekly Feed - 7 hours 9 min ago
IT services firm will provide London police force with IT services for six years as part of new deal
Categories: Computer Weekly

True digital sustainability requires addressing software waste. Refactoring bloated code cuts computational overhead, lowering enterprise energy consumption and hardware churn

Computer Weekly Feed - 7 hours 9 min ago
True digital sustainability requires addressing software waste. Refactoring bloated code cuts computational overhead, lowering enterprise energy consumption and hardware churn
Categories: Computer Weekly

A malfunctioning electronic visa status and administrative errors on the part of the Home Office have left a single mother and her son facing removal from the UK, despite her proactive efforts to legally secure a new visa

Computer Weekly Feed - 7 hours 9 min ago
A malfunctioning electronic visa status and administrative errors on the part of the Home Office have left a single mother and her son facing removal from the UK, despite her proactive efforts to legally secure a new visa
Categories: Computer Weekly

Local councillors welcome the prospect of becoming one of the world’s largest centres for AI computing, but a spurned rival thinks they cheated and still got a bad deal

Computer Weekly Feed - 7 hours 9 min ago
Local councillors welcome the prospect of becoming one of the world’s largest centres for AI computing, but a spurned rival thinks they cheated and still got a bad deal
Categories: Computer Weekly

The transfer of the Civil Service Pension Scheme to a new administrator has been a disaster, but why?

Computer Weekly Feed - 7 hours 9 min ago
The transfer of the Civil Service Pension Scheme to a new administrator has been a disaster, but why?
Categories: Computer Weekly

Pages