Hacker Newsnew | past | comments | ask | show | jobs | submit | bestcommentslogin
Most-upvoted comments of the last 48 hours. You can change the number of hours like this: bestcomments?h=24.

> Claude Opus 5.5 is our first release since we called for pacing the frontier.

Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.


> It found the U.S. “failed in its obligation to do everything feasible to verify” that the school was a military objective and that the failure “went beyond mere negligence.” The report said the United States “directed the strikes at the building of the school while being aware of a substantial risk of striking a civilian object and acting recklessly as regards the possibility that this would happen.”

Reading the details, "AI" doesn't really seem like the culprit -it's a scapegoat.

The intelligence that it was no longer a military target never entered the target database, the team that was responsible for vetting the target list was gutted, and said team was never even consulted.

The White House wanted 1000 targets and pulled from their database without any due diligence. Whether it was an AI call or an SQL query - this was from pure human maliciousness and incompetence.


As I have been saying for years:

Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are. If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer. You might as well transfer those bits to me directly, rather than having the LLM add on an extra superfluous 700 bits that I then have to filter out.


Here's my constant question:

Everyone's going so fast that they keep hitting walls. Review, CI, product asking for things, whatever.

Why have we not seen an improvements in products?

While every post and thread feels like a 90's wall street office, the new android and iphone ship with fewer features than usual. No indie guys come up with a linux-sized alternative OS. Switch 2 remains unhacked. Windows takes 3 seconds to show the right click menu.

Is everyone just running full speed in circles or something?


I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

If you’re releasing an open model going forward, please consider offering the community more of this transparency!


Sorry folks, this is a rollout artifact, we needed a way to turn this off remotely via feature flags if it broke something, and with telemetry off you don't get those. It's already been fixed as part of v2.1.281 releasing today.

The mod is source available here: https://github.com/anthropics/claude-code/tree/main/mods/age...

Apologies again folks, this was a fully human error on my part - I should've found a better way to launch with a kill-switch.


Not trying to start flamewars, but more and more I wonder how much are people willing to put up with such practices. On linux for 15+ years and everytime I try to use mac or win, it is an ordeal. And ads. F*cking ads in a system someone has purchased. And spying. Really user hostile environment. No, thank you.

GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

Here's GPT-6 Luna pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

And GPT-6 Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Scroll to the bottom for the GPT-6 Sol max one: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

For comparison, here are the pelicans I got for GPT-6 Astra: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...

The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.


At this point, no one seems capable of keeping a large database safe. I assume all medical and biographical information that exists is in the hands of the major state actors.

China hacked 22.1 million records of US government employees:

https://en.wikipedia.org/wiki/2015_Office_of_Personnel_Manag...


“Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself

Finally that price drop

   Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25

Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.

If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor


I doubt John Ternus will change direction anytime soon, since it will look like he’s reversing the plethora of eyesore ads that Tim Cook added over his tenure.

But the Apple with ads is not the Apple that had some taste and discernment in the past. For a long time I’ve visited the App Store’s app update page directly (tap and hold on App Store icon to see the context menu option). Anytime I inadvertently go to the App Store home page or the few times I search, it’s an ad filled disaster!

From this article

> repeatedly attempting to prod customers towards even more of the company’s products might seem cheap, even distasteful.

From a recent post by John Gruber:

> Steve Jobs in 2011: 'We Build Products That We Want for Ourselves, Too, and We Just Don't Want Ads' [1]

Looks like Tim Cook, John Ternus and Eddy Cue really enjoy being swamped with ads in their products. Will there soon be a time when Apple executives start carrying some other brand’s devices with them to avoid having a rotten experience?

[1]: We don’t want ads https://daringfireball.net/linked/2026/07/28/jobs-we-dont-wa...


It happens. There use to be a joke during the first big DC build out phase that went like if you're ever going into the wilderness take a 1ft length of fiber optic cable with you. If you get lost bury it and a back hoe operator will appear and sever it within an hour. You can get a ride back with them.

> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.


I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.

The meaning isn't clear at all. So open for interpretation that it is meaningless. That's the whole fucking point. For all I know they are "pacing the frontier", or not. The fact that there's no meaning to it let's you know that it was a pointless waste of tokens and attention.

Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.


This article only obliquely mentions in the last paragraphs, but there's a parallel Chinese program that recently brought back samples from the moon, and will attempt Mars sample return, launching in 2028,

https://en.wikipedia.org/wiki/Tianwen-3


Because Mac laptops are lightyears ahead of any other hardware.

Build quality (aluminum), charging (magsafe), screen resolution, battery life, noise (or lack thereof), trackpad…

It’s like other vendors aren’t even trying.

Also the OS mostly “just works”, just last week my Linux laptop disabled NVIDIA GPU (and almost bricked itself?? Not sure, had to fix apt) during automated updates


Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.


A good chunk of what my company has been doing with AI falls into either burning down our known tech-debt and "easy wins" that no one ever had the bandwidth to approach... And improving / automating our processes. The former is having a direct and meaningful impact on the quality and availability of our services.

Our QA, formerly a fairly frequent blocker of all our releases, are doing more in-depth reviews and catching issues earlier in our release process. They have become unblocked to the point they are actively chasing down work that starts to slip.

We have cleaned up and tuned both our security alerts and operations logs and improved our tenant isolation in our service in a way that makes customer and formal audits SIGNIFICANTLY easier.

We're setting ourselves up for faster human development of the hard-things. Our development environment and infrastructure are faster, cleaner, more auditable processes, and cheaper overall to operate.

These fixes mostly don't show up in our product change logs, and definitely don't fall into "new features". It would largely be invisible to the outside world, but our costs are going down (though to be fair, not offsetting the spend on AI to date), internal productivity has improved, operational incidents are down, and customer satisfaction is up.


Anyone wanting to be involved in the future of scientific research should probably start learning Mandarin. We're going the way of the Soviets, where political officers will quash any science that doesn't fit their ideology. https://arstechnica.com/science/2026/09/trump-planning-to-ha...

> About 20 Markdown files described browser use, connectors, payments, credentials, data handling, generated files, voice, goals, and scheduling.

This the state of software engineering in 2026.

Edit: clarified engineering to software engineering, which is more correct


We ended up with a Samsung fridge in our new home. It has no external display or any other indication that it is a "smart" fridge. I only found out when a friend visited who has a Samsung phone and the phone offered to connect to the fridge. We spent the next 30 minutes removing panels from the fridge and found the wifi/bluetooth antenna on a small board under the top-right hinge cover. The board was disconnected and the fridge continues to operate normally.

I am now eyeing my dishwasher very suspiciously.


The problem imo is the slow deterioration of institutional knowledge that offloading the mental task of wisdom gathering to AI is causing.

One interesting comparison is to the history of manufacturing. West/America decided one day that manufacturing would be cheaper to outsource and better (short term) profit was to be made by outsourcing it all to China. The institutional expertise started to deteriorate, to the point that America simply didn't even have the capacity, or expertise anymore to produce stuff (such as grill brush [1])

I feel like you could take all the handwavy comment that are made today to dismiss this caution, and find equal dismissal back then when companies were actively outsourcing the manufacturing.

"I'm coding 10x faster" "look at the output velocity per employee"

"we are producing much more (in China)" "look at profit / number of (manufacturing) employers"

Seems ok if you're American / Chinese but I'm struggling to understand how the rest can be OK with allowing institutional knowledge to deteriorate while having an active dependency to the former two. We already see this with the tech dependency towards USA and manufacturing competition from China.

[1] https://youtu.be/3ZTGwcHQfLY


It's funny, i push back on pull requests because there is too much description now - a 20 line change has pages and pages of generated description, rationalisation for why it is safe, defense of each design decision, analysis of risks and side effects. People are indignant, you're rejecting my change because there is too much documentation? And my response is, I don't have time to read it and you put me in the position where I can't afford not to - because approving the PR implies I did and accepted it. The investment to read all that for the value of a code change that I'm one prompt away from doing myself if I cared is just not high enough. So it's rejected.

The Mosaic browser (1993) had full text history search.

We then got a bookmark system that was every bit as terrible as a web directory.

It stayed that way. The delicious search revenue made organizing websites uninteresting. That obscure website you enjoyed a decade ago but don't even remember, they had lots of traffic like you. No point updating or keeping it online. You can't have rss in Firefox but here is a Facebook like button in your address bar in stead.

I've tried to maintain the bookmark menu but I rarely use it since everything is dead. Why aren't browsers storing a text version of the bookmark? Did people in 1993 have more resources than us? Should I be afraid it grows to a few GB over the decades?

A good few dead websites have a backup some place but there is no automation to find it. If you had a string of text from a page you might be able to search for it. If the page found is highly similar we might automate the process to have alternative location for bookmarks with a nice warning dialog.


I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.

Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.


The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.

AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.


Anyone else more excited about Chinese models than American models these days? Big thing for me is affordability.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: