Feed aggregator
SF $192 fine for teaching tennis to his own children
Article URL: https://twitter.com/Polymarket/status/2084790486590517503
Comments URL: https://news.ycombinator.com/item?id=49967465
Points: 1
# Comments: 0
Show HN: Jiwo - Small decision models topping decision index leaderboard
Spent the weekend turning small LLMs into decision models, with some good results and a lot of learnings along the way. Topping the decision index leaderbard for their categories https://huggingface.co/spaces/multimodalart/jev-decision-ind...
Sharing two of them: a sub-1B model and a 4B model, both Qwen3.5 fine-tunes. When evaluated on the decision Index shared last week, the 0.8b tops the sub-1B category, 45.7% above the best other Qwen3.5-0.8B fine-tune on the leaderboard. and the 4b comes in second in the 3–6B class.
Most of work was data calibration, finding external datasets and readapting them towards this scenario, so i expect a lot of improvements and work like this to come from community and push these numbers even higher. Even more if we get qwen 3.8 releases for these model categories.
Comments URL: https://news.ycombinator.com/item?id=49967455
Points: 1
# Comments: 0
Green party votes to adopt policy that equates Zionism with racism
Article URL: https://www.theguardian.com/politics/2026/oct/04/green-party-votes-adopt-policy-zionism-racism
Comments URL: https://news.ycombinator.com/item?id=49967447
Points: 3
# Comments: 0
US closely monitoring case of lab worker who possibly died of plague in Siberia
Thoughts on AI
Article URL: https://write.as/davepolaschek/thoughts-on-ai
Comments URL: https://news.ycombinator.com/item?id=49967439
Points: 2
# Comments: 0
Benchmark in Milliseconds
Article URL: https://matklad.github.io/2026/10/05/benchmark-milliseconds.html
Comments URL: https://news.ycombinator.com/item?id=49967427
Points: 1
# Comments: 0
Interfaze-1-lite: the first open-weight model for deterministic task
Article URL: https://huggingface.co/interfaze-ai/interfaze-1-lite
Comments URL: https://news.ycombinator.com/item?id=49967425
Points: 2
# Comments: 0
Show HN: Self-bench – benchmark coding agents on real-world software
Hey HN,
Today, we're launching selfbench.dev, an open-source tool that lets you create and run evals automatically from your PRs.
Every benchmark with sufficient trust eventually gets benchmaxxed (Goodhart's law) - the labs are incentivized to maximize their scores on that benchmark, which isn't predictive on whether it'll actually work within your setup. This has been a time-consuming process that only the largest companies can afford to do, so we built self-bench to fix this.
Self-bench uses agents to author / review Harbor environments generated from your PRs. You then approve every eval that the agent created, and can model/harness evals concurrently on sandboxes.
To save you money, we also allow you to connect your OpenAI and Claude subscriptions so you don't have to pay raw token costs.
We ran evals on several large codebases like
- Posthog (https://selfbench.dev/PostHog/posthog)
- Next.js (https://selfbench.dev/vercel/next.js)
- Sentry (https://selfbench.dev/getsentry/sentry)
- Pi (https://selfbench.dev/earendil-works/pi)
- as well as other fantastic open-source projects (all on selfbench.dev)
From our evals:
- Kimi K3 and GLM 5.3 are almost always more expensive, yet less performant than models like GPT-6.1 Sol and Claude Opus 5.5, due to token efficiency.
- GPT-6 Luna is almost always the most effective "cost-efficient" model we've tested, not open-source models.
We'd love for you to try this on your codebase and let us know what you think!
Comments URL: https://news.ycombinator.com/item?id=49967408
Points: 2
# Comments: 0
BrontoDB: The Polymorphic Database for Observability Data
Article URL: https://bronto.io/blog/brontodb-the-polymorphic-database-for-observability
Comments URL: https://news.ycombinator.com/item?id=49967402
Points: 2
# Comments: 0
Talk of an AI kill switch abounds, but what would such a thing actually look like in practice, and what would it mean to ‘press’ it?
IT services firm will provide London police force with IT services for six years as part of new deal
True digital sustainability requires addressing software waste. Refactoring bloated code cuts computational overhead, lowering enterprise energy consumption and hardware churn
A malfunctioning electronic visa status and administrative errors on the part of the Home Office have left a single mother and her son facing removal from the UK, despite her proactive efforts to legally secure a new visa
Local councillors welcome the prospect of becoming one of the world’s largest centres for AI computing, but a spurned rival thinks they cheated and still got a bad deal
The transfer of the Civil Service Pension Scheme to a new administrator has been a disaster, but why?
Flock Blamed for Wrongful 13-Day Imprisonment and a Police Stalking Incident in Florida
People are asking ChatGPT to help them decide how to vote in the midterms
Article URL: https://www.npr.org/2026/10/05/nx-s1-5977852/ai-chatbots-midterm-election
Comments URL: https://news.ycombinator.com/item?id=49965833
Points: 1
# Comments: 0
I built a $5/month TV app because I got tired of how expensive streaming is
Article URL: https://streamx1.com/Phone.apk
Comments URL: https://news.ycombinator.com/item?id=49965821
Points: 1
# Comments: 0
Meta Rushed to Fix Muse 'VM Escape' Vulnerability Soon Before Launch
Article URL: https://www.404media.co/meta-rushed-to-fix-muse-vm-escape-vulnerability-immediately-before-launch/
Comments URL: https://news.ycombinator.com/item?id=49965808
Points: 1
# Comments: 0
