Hacker News

Horror Technologies

Hacker News - 6 hours 21 min ago

Imagine if there was a technology developed where the American government could add other people’s minds to your bodies. This technology could have civilian and non-civilian use cases. One could also run miniature AI models in such a way that it runs inside your head entirely.

Comments URL: https://news.ycombinator.com/item?id=49967547

Points: 1

# Comments: 0

Categories: Hacker News

Show HN: Jiwo - Small decision models topping decision index leaderboard

Hacker News - 6 hours 29 min ago

Spent the weekend turning small LLMs into decision models, with some good results and a lot of learnings along the way. Topping the decision index leaderbard for their categories https://huggingface.co/spaces/multimodalart/jev-decision-ind...

Sharing two of them: a sub-1B model and a 4B model, both Qwen3.5 fine-tunes. When evaluated on the decision Index shared last week, the 0.8b tops the sub-1B category, 45.7% above the best other Qwen3.5-0.8B fine-tune on the leaderboard. and the 4b comes in second in the 3–6B class.

Most of work was data calibration, finding external datasets and readapting them towards this scenario, so i expect a lot of improvements and work like this to come from community and push these numbers even higher. Even more if we get qwen 3.8 releases for these model categories.

Comments URL: https://news.ycombinator.com/item?id=49967455

Points: 1

# Comments: 0

Categories: Hacker News

Thoughts on AI

Hacker News - 6 hours 30 min ago
Categories: Hacker News

Show HN: Self-bench – benchmark coding agents on real-world software

Hacker News - 6 hours 32 min ago

Hey HN,

Today, we're launching selfbench.dev, an open-source tool that lets you create and run evals automatically from your PRs.

Every benchmark with sufficient trust eventually gets benchmaxxed (Goodhart's law) - the labs are incentivized to maximize their scores on that benchmark, which isn't predictive on whether it'll actually work within your setup. This has been a time-consuming process that only the largest companies can afford to do, so we built self-bench to fix this.

Self-bench uses agents to author / review Harbor environments generated from your PRs. You then approve every eval that the agent created, and can model/harness evals concurrently on sandboxes.

To save you money, we also allow you to connect your OpenAI and Claude subscriptions so you don't have to pay raw token costs.

We ran evals on several large codebases like

- Posthog (https://selfbench.dev/PostHog/posthog)

- Next.js (https://selfbench.dev/vercel/next.js)

- Sentry (https://selfbench.dev/getsentry/sentry)

- Pi (https://selfbench.dev/earendil-works/pi)

- as well as other fantastic open-source projects (all on selfbench.dev)

From our evals:

- Kimi K3 and GLM 5.3 are almost always more expensive, yet less performant than models like GPT-6.1 Sol and Claude Opus 5.5, due to token efficiency.

- GPT-6 Luna is almost always the most effective "cost-efficient" model we've tested, not open-source models.

We'd love for you to try this on your codebase and let us know what you think!

Comments URL: https://news.ycombinator.com/item?id=49967408

Points: 2

# Comments: 0

Categories: Hacker News

Pages