Hacker Newsnew | past | comments | ask | show | jobs | submit | Silhouette's commentslogin

I feel like every time this conversation comes up now someone has to remind everyone that an LLM is just a mathematical model. An LLM can't do anything except produce a stream of output tokens. The problems we keep seeing are tools that interpret those output tokens as actionable instructions without an adequate framework and safeguards for how they operate.

Data from LLMs being processed by these tools should be treated the same as any external input into any software system: parse - don't validate - to convert to a systematic representation with deterministic consequences and then consider those consequences within a clearly defined and limited framework. You never trust data from external sources verbatim. And you never try to use vague human language when you need to describe precise technical details unambiguously.

We learned these lessons a very long time ago in programming. It's why we have programming languages in the first place among countless other examples. But way too many people are so infatuated with LLMs and agents that they've already forgotten the basic principles of their craft after only a few months.


> LLM is just a mathematical model

Yes those of us who bothered to know the internals know of this. But the marketing says that these are magic tools.. So that's gotta be a shock for them, but the joke is the people who irresponsibly use this won't ever read this!


The thing I find concerning lately is that even a lot of technical people seem to be jumping on the hype train this time around. Obviously LLMs have become very useful tools for assisting some technical tasks but even SOTA models are nowhere near reliable and predictable enough to trust their output completely as YOLO mode agentic workflows effectively do.

Given the nature of LLMs I don't think they can ever clear that bar without some other element being introduced. The nondeterminism and chaotic nature of LLM output is enough to rule them out as a reasonable foundation for any fully automated system that would be controlling anything potentially dangerous or damaging.

But it seems to be heresy at the moment to even suggest that the future might not be bright if everyone just relies on agents driving LLMs to do all the real work. The number of people I've encountered in the past year who I'm fairly sure are smart and technically capable and yet who are also now happy to do development and other tasks either without any human in the loop at all or with at best a cursory LGTM level review before approving the LLM's output is remarkable.


> The nondeterminism and chaotic nature of LLM output is enough to rule them out as a reasonable foundation for any fully automated system that would be controlling anything potentially dangerous or damaging.

You could say the same for humans. The difference is that humans have been conditioned to be extra cautious about things that could get them fired, and there is no benchmark for Meta's model developers to benchmax about that.


Agreed. I see people using LLM to reply to slack or to commit a git branch!!

Execs expecting to double your workload just because you have LLMs. It's insane at this point


Option A: If you have a tiger in your room then make sure it's properly caged. Put up warning signs so everyone knows. Add physical barriers to prevent people too young or impaired to read the warning signs from approaching close enough for the tiger to reach out of the cage and maul them. Ensure adequate processes are in place for feeding the tiger at regular intervals using a safe method and clearing out the mess from the cage. Provide noise protection for everyone in the building so they don't get freaked out when the tiger complains vocally about its situation. Take into account evolving animal rights legislation and ensure adequate processes are in place for the tiger to exercise freely in a large open space. This open space will also need to be protected by safety barriers and warning signs as well as supervised by trained operators able to contain a wild tiger if it gets loose and deal with any injuries or damage it causes. Budget for all of this and ensure there is a long term plan for maintaining the tiger and everything that goes with it.

Option B: Do not put a tiger in your room.


Option A: If you have 1 ton metal machines powered by exploding liquid driving at 60 miles per hour around your neighborhood make sure the wheel is properly pointed in the right direction. Designate specific areas where the machine are allowed to move. Train parents to keep their children or impaired away from those areas. Ensure adequate processes are in place from training the drivers of the machines so they know where they are allowed to drive and where not. Take into account evolving environmental regulations. Ensure safety barriers are used in areas where the vehicles are liable to lose control. Place signs around to remind forgetful drivers of the specific rules of driving in a specific area, or issues such as ice and snow that could cause the vehicle to crash. Make sure drivers pay large insurance premiums to ensure they can pay for any damages caused by their vehicle. Give traffic police their own vehicles in order to apprehend anyone not following the rules. Budget for all of this and ensure there is a long term plan for maintaining the car and everything that goes with it.

Option B: Walk everywhere


Option C: Let the tiger drive, and pretend there is no human to hold accountable

Given the number of places in the world that have been severely restricting or banning car use to promote walking and other safer and more environmentally friendly alternatives I'm not sure your analogy is making exactly the point you intended.

Option C: put the tiger in someone else's room :)

Option D: Be the tiger, my friend :)

instructions unclear; tried putting tiger into a cup, then a bottle

now all my flows crash :D


One significant problem with the TV licence here in the UK is that the people who end up paying it and the people who actually use the services it pays for are now very different groups. The BBC is heavily funded by TV licence fee payers but also provides numerous radio and online services. Many of those are available even to those without a TV licence. Meanwhile people who use few or no BBC services may still have to pay for a TV licence because they watch something else live but the money they're paying is mostly going to be used to support the BBC.

Also it's a silly little tax that raises relatively little money, incurs significant overheads in collection, suffers from significant non-payment, and has an aggressive enforcement operation that should clearly be illegal itself for the way it behaves.


That's pretty much the situation in the UK if you are a sole trader. If you set up a limited liability company, that's a separate entity and has its own tax reporting requirement, but nothing onerous.

It's becoming more onerous all the time to run a small business in the UK though.

We just got charged £100 just to essentially tick one box with Companies House saying "nothing changed since last year".

The main tax paperwork is the annual financial statements. These are getting more demanding over the next couple of years with the simplified versions that small companies have been allowed to file looking like they will no longer be acceptable after that time.

You now have to file tax returns like PAYE and VAT using approved software. Most of that software - even the big names - ranges from OK but with significant problems and limitations to simply awful.

I heard recently that personal self assessment tax returns for company directors are about to get more complicated as well.

If you have business premises then you have to deal with business rates. Often these are still based on valuations from a time when dinosaurs roamed the planet. Some small businesses have already closed down due to this becoming unaffordable.

I've lost count of how many different kinds of insurance we now have. All have a cost of course.

If you're working solo or the only person from your company working with a specific customer/client then you need to worry about Off Payroll Working as well. That's still a huge mess of ambiguity decades after IR35 was introduced and likely incurs further expenses for legal reviews and insurance policies that shouldn't even need to exist.

Control any personal data? Make sure you register with the ICO and pay their fee too.

Sell online or even run a website? Make sure you know the requirements for information you have to provide and it's in exactly the right format for compliance. (I've even seen businesses working in the compliance industry failing on this one and risking regulatory sanctions.)

The public sector is increasingly pushing for things like Cyber Essentials if you want to be part of its supply chain. That is an absolutely textbook example of regulations that were clearly written for large businesses using a certain operating system. The only options are 100% compliance or failing and 100% compliance is tricky for sole traders or microbusinesses to achieve without disproportionate expense and administration. If you need the certified version then that's another cost and the disruption of being audited to deal with.

We haven't even talked about everything you have to do if you want to take on employee #1 who isn't one of the founders/owners yet.

I don't think it's realistic to say that the UK is supportive of small businesses today. It never really was but at least the the potential financial benefits used to justify the overheads. Now the taxes are far higher - you end up paying double-digit higher overall rates if you set up a Ltd that can actually grow as a business compared to simply trading as a self-employed sole trader now - and the overheads and costs are also far higher in numerous ways. It's no surprise that people are deterred from going into business here and that's really bad for the country because we need the people willing to take the risks and put in the effort to build the next generation of businesses to support our economy and provide much-needed new jobs.


Anecdotally my own experience is consistent with this with the same platform+browser combination.

It appears that there are now significant numbers of sites - or at least noticeable parts or features of sites - that rely on Google-specific APIs and haven't been tested on other browsers.

However visiting those sites from Apple devices is often similarly frustrating. I'm not sure this is an anti-Firefox thing. It seems more of a not realising there are other browsers apart from Chromium-based ones thing.


I kinda wondered myself if this was some kind of coordinated effort to make Firefox not work. It wouldn't surprise me if it was later discovered that Google paid some Linux subsystem maintainer to slip in some nefarious code somewhere that altered the way Linux implements rendering specifications ever so slightly such that Firefox appeared broken. Just the conspiracy theorist in me.


I appreciate your positivity. However it is difficult to reconcile your argument with the reality that executives at tech companies frequently behave in ways that satisfy the investors but clearly do not produce good results for the people using their technology. Indeed in numerous cases today the technology is openly hostile to its own users' interests. The regulations are there to protect the users (or customers or clients or patients). If anything they are there to prevent protecting the profit margins of the businesses - because hiring people to pump up the profits is often in conflict with the other objective.


What would credentialing software engineers do to curtail leadership incentive problems?

Hardware technology jobs have some of these credentials and are part of that system.


What would credentialing software engineers do to curtail leadership incentive problems?

The same as in other regulated professions. It would give the people on the front line who can see the consequences of the corner cutting an effective right to say "No, we're not doing this user hostile thing". This would be a significant barrier because it wouldn't be legal to ship software without the required professional approval and the professionals would be heavily incentivised not to sign off any corner cutting because they would be personally and professionally responsible for any adverse consequences if they did.


That’s not how these systems work in reality. I’m actually an extreme pessimist.

What actually happens is that you as an engineer become paid to frame the desired leadership goals in compliant terms.

It’s similar to how lawyers and regulatory compliance rarely change the product. They change how the product is talked about and framed for the purpose of regulation.

This skill among engineers is highly sought after in large companies with political organizations and greater encouraging it changes the composition of the workforce.

Top civil engineers don’t design buildings.


In the physical engineering teams I've worked with the culture you just described is not what happens at all. There are always some people who seemingly live to circumvent the intention of rules lurking around the edges of regulated industries. But my own experience has been that real engineers very much value real engineering - it's often why they got into the field in the first place - and will push back hard against and if necessary refuse to sign off anything they consider inappropriate for the job. They won't be impressed at all by someone's big title and even bigger budget while they're exercising that professional judgement either. Some of these people have had pretty stellar careers so I think it's safe to say their professional conduct hasn't damaged them with their employers or clients.


> circumvent the intention of rules lurking around the edges of regulated industries

That’s not what I said,

For every human activity there is an underlying reality and there is a social component of how it’s framed or talked about.

Regulatory compliance is 90% social and 10% reality.

So a focus on compliance means engineers spend less of their time on reality.

> They won't be impressed at all by someone's big title and even bigger budget while they're exercising that professional judgement either

Correct. But they do care about their management chain.

Imagine a new grad telling their boss “we aren’t going to build it like that because I learned X in school”. They have the same credentials!


Regulatory compliance is 90% social and 10% reality.

Again our experiences are on opposite ends of the spectrum. For safety issues in particular getting an engineer to sign off some plan when they will be accountable for that authority later is often easiest if you simply design the thing properly.

I have certainly seen compliance become a box-ticking exercise in other contexts but usually this seems to happen when the professionals involved were not personally responsible for their own decisions.

Imagine a new grad telling their boss “we aren’t going to build it like that because I learned X in school”. They have the same credentials!

We appear to live in different realities on this one too. In my reality new grads are not the people signing off major decisions in regulated industries and no-one would seriously suggest that a new grad's degree was an equivalent credential to the years of demonstrable professional experience and peer review that are typically required to reach a level of professional qualification where someone does have the authority to sign off those big decisions. Getting an undergraduate degree in a subject like engineering or medicine or law is just a foot in the door. The real work starts afterwards.


I think we are in full agreement!

> For safety issues in particular getting an engineer to sign off some plan when they will be accountable for that authority later is often easiest if you simply design the thing properly.

That's exactly right. An engineer aims to build a safe and reliable product (not because the credentials tell him to)! The regulatory myth is that all products are death machines until being redeemed.

That's why I say it's 90% social. The engineer builds a reasonable product. With regulation they do the same, but now they need to do work to frame that same work in regulatory terms.

Real fixes are made. The value isn't 0, but it comes at that cost.

> no-one would seriously suggest that a new grad's degree was an equivalent credential to the years of demonstrable professional experience

That's exactly what I'm saying! Professional experience and reputation within the field is what:

1. gives someone influence and credibility. 2. results in safe and reliable engineering projects.

And note that those are actually informal defined qualifications. It's NOT the credential!

So trying to add credentials to software is an attempt to bump the quality of new grads and has little to no impact on the quality of engineering leadership. As you said, the recent undergrads already lack power and influence, and the credential is simply the bare minimum to participate.


> They won't be impressed at all by someone's big title and even bigger budget while they're exercising that professional judgement either. Some of these people have had pretty stellar careers so I think it's safe to say their professional conduct hasn't damaged them with their employers or clients.

It really depends on the status of the profession in society and the company. I'd expect lots of software organisations to fire a bunch of people looking for people pleasers until this becomes an accepted part of the business approach.


I'd expect lots of software organisations to fire a bunch of people looking for people pleasers until this becomes an accepted part of the business approach.

No doubt. Imposing this kind of professional standards to regulate an industry as rich as tech would never work unless the penalties for cutting corners involved making the offending organisations significantly less rich very quickly. They would need to be taught a very clear lesson that hiring people pleasers had become an expensive mistake.

As I commented elsewhere - the problem then becomes who gets to define what the proper path is. For example destroying companies because they chose not to follow the latest sage advice from anyone who once signed the Agile Manifesto does not seem like a good way to promote better quality software to me. And yet it seems highly likely that those are the kinds of people who would initially be engaged as "experts" by those seeking to establish the regulatory environment.

I don't want people like them. I want the quiet, unassuming developer you've never heard of because they're the principal engineer of a team you've also never heard of that has been developing life saving medical equipment without a single significant failure in a live environment for the 15 years since their first device went into use at a local hospital. Get me those people to write the rules - starting with what is acceptable practice when developing software that really needs to work and letting the people who know how to achieve the most challenging results figure out how to tone everything down for applications where imperfections might be more acceptable - and then we can talk about whether regulating software development effectively is now a viable proposition.


> As I commented elsewhere - the problem then becomes who gets to define what the proper path is. For example destroying companies because they chose not to follow the latest sage advice from anyone who once signed the Agile Manifesto does not seem like a good way to promote better quality software to me. And yet it seems highly likely that those are the kinds of people who would initially be engaged as "experts" by those seeking to establish the regulatory environment.

This is why any such regulation is likely to end up in a much better state if it's driven by actual practitioners. However, given how many software people are wildly against this, it seems unlikely to happen and so we'll end up in the less good state you note above.


It looks like we agree here. I too think any useful regulation should be specified primarily by experienced practitioners who have been achieving demonstrably good outcomes. And I too think that in reality it would mostly likely be a very different type of person who got to write the rules - which is why I don't think our industry is ready for that kind of regulation and I believe introducing it now would be counterproductive.


> It’s similar to how lawyers and regulatory compliance rarely change the product. They change how the product is talked about and framed for the purpose of regulation.

This is true, and basically what tends to happen is that if the Head of some compliance function (e.g. internal audit) is causing problems for the business, then they are replaced with someone who won't cause such problems.

It's still better than nothing. Like, software basically runs our society now, so either software professionals get together on this, or regulations will be imposed on us, and they will be much worse than what we'd get in the first option.


> It's still better than nothing.

Once again the alternative is not nothing. The most important factor is that they are stakeholders in a project with influence. That is the reality right now, even without credentials.


> Once again the alternative is not nothing.

Can you help me understand what the alternative is?


Yes the framing above is “business people don’t listen to engineers but if they had a credential they would legally be forced to”.

And I think both are false.

The top engineers on a project are collaborators with leaders in other areas like marketing, sales, IT, legal, etc. And all those have influence on a project. A business person who says “fuck what my engineer says” is not a good leader and won’t have that group’s trust or support. They all want to work together.

So by that process engineering has a seat of influence.

That exists without credentials, and credentials are not what gets you in that seat. Reputation and experience are.

Business leaders don’t do everything engineering says. They also don’t do everything the lawyers say! And having an additional legal backing would change some of these engineering conversation, but not fundamentally.


More credentialing and requirements would do plenty to curtail leadership incentive problems. First, as the sibling points out, it gives the licensed engineer a solid ground to stand on when refusing a stupid leadership order — "I'm not going to lose my license for that stupid idea", and a serious incentive to do so, as well as solid job prospects if he does get canned for it (because there's a limited pool of credentialed engineers).

You glibly say "More schooling does not mean more trustworthy or careful doctors.". Yet the schooling and credentialing clearly cuts off huge numbers of would-be doctors who never pass the exams, never graduate med school, or never even get into med school, or decide it is too difficult in the first place. In the software realm, those people just go to some boot camp and they're off to the races...

Another huge aspect of credentialing is required ongoing education, which REQUIRES physicians and engineers to take updated continuing education just to maintain their license. This again continuously improves the talent pool.

And, if your main concern is that they be "more trustworthy or careful", credentialing also helps that by finding the worst, least trustworthy and careful and cancelling their license, so they are NOT doctors anymore. The untrustworthy or careless SWE just gets a new job to ruin stuff elsewhere, probably taking one from the actual good engineer because their talent is not engineering, but bullshitting.


> it gives the licensed engineer a solid ground to stand on when refusing a stupid leadership order

I disagree to the extent to which this is real leverage. It changes the language and approach, but not the outcome, Ask a civil engineer the degree to which they can fight their leadership on these grounds.

> clearly cuts off huge numbers of would-be doctors who never pass the exams

Yes the fallacy is that more exclusive is better. You don’t understand the traits you select for.

> those people just go to some boot camp and they're off to the races

I don’t see any kids who just got off a boot camp running large software projects. Does this happen at your workplace? Why not?

> This again continuously improves the talent pool.

I just disagree to the extent to which the talent actually increases.

The people who excel already learn and study all time.

This slightly raises the floor by forcing the least curious person to be exposed to some PowerPoints and videos.

> The untrustworthy or careless SWE just gets a new job to ruin stuff elsewhere

They can only ruin the extent of responsibility and scope given to a new hire with no reputation.

> least trustworthy and careful and cancelling their license

Once again you assume the system works as stated. I think it actually selects against those who are bad at avoiding responsibility and not legally savvy.

The image that comes to mind is someone who made a mistake, cares a ton about medicine, and hates the organizational administration.


Wow you have a relentlessly unrealistic view of things.

>>Ask a civil engineer the degree to which they can fight their leadership on these grounds. Both civil engineers and doctors both can and absolutely do refuse to sign off on unsafe situations.

That does not mean they detect them 100% of the time, or never cave to pressure, but they absolutely do. We just never hear about the incidents that didn't happen because a doctor or engineer forced the right solution or no action — precisely because nothing newsworthy happened. We DO hear about the ones that did happen, Therac-25, Mars Climate Orbiter loss, Cloudflare outage, it is endless

>>I think it actually selects against those who are bad at avoiding responsibility and not legally savvy.

You might think that, but clearly you have never read even the summaries of cases where doctors lost their licenses. Hint: it was not some mistake that could have been covered up by better schmoozing. If anything, the system is too lenient.

The rest isn't even worth the bytes to respond; just handwaving an attitude. There are very good arguments to not have licensing on software engineering but you are not making them


> Wow you have a relentlessly unrealistic view of things.

I noticed from your response that you didn’t really refute the claims, but that suggesting that systems don’t achieve their stated goal gives you a distasteful feeling.

> We just never hear about the incidents that didn't happen because a doctor or engineer forced the right solution or no action

And the same is true of software. I and my peers tell my bosses ideas are bad all the time. We don’t need a credential to do that. And the credential is not what gave us trust with that decision maker.

> If anything, the system is too lenient.

Correct. Bad doctors continue to keep their jobs all their time.

So that’s my point. What is the criteria that distinguishes those cases? Both doctors made a medical error. Which one gets off and which one gets fired? The doctor who is more focused on medicine is likely the one less skilled at navigating the legal problem.

The doctors making mistakes and keeping their jobs are a pathological minority that is reinforced by their credential giving them authority to operate.


The problem with making programming a regulated profession is who gets to define the regulations. You are probably imagining that it would be expert developers with a track record of success. I would bet on it being the people who write lots of blog posts and books about programming and give lots of conference keynotes - whether or not those people have any evidence base to support their advocated policies or any personal track record of delivering good software.


> You are probably imagining that it would be expert developers with a track record of success.

I'll admit I am unsure how such individuals would be chosen. I imagine it would be prudent to learn what processes are used by other licensed professions when choosing such individuals.

> who write lots of blog posts and books about programming and give lots of conference keynotes

I do not think people who write blog posts and give conference keynotes should be awarded with roles that regulate the profession because they write blog posts and give keynotes. I would hope that any regulation is evidence driven.


> I'll admit I am unsure how such individuals would be chosen. I imagine it would be prudent to learn what processes are used by other licensed professions when choosing such individuals.

Licensing organizations disproportionately attract people who enjoy "administrating" over "doing."


Other those who prioritize profits over providing good customer service.


People who enjoy administrating rarely think you’re the “customer” lol.


> I would hope that any regulation is evidence driven.

Have you found existing regulation in any field that's evidence-driven?

My impression is that regulation is not made in a way that has much to do with evidence, but I would be delighted to be shown evidence I'm wrong.


Aviation comes to mind. Most aviation regulations exist because people died.

It's hardly a perfect system, but the problem is mostly that it's too conservative and risk aware, making it difficult to impossible to innovate and ignoring the risk this creates. (The world's most popular light aircraft is the 1950s-vintage Cessna 172, mostly because it's impossibly slow and costly to get a reasonably priced modern competitor certified.)


Thank you for the data point. I guess I had overlooked that.


Have you never heard the adage that regulations are written in blood?

Plenty of regulations in many many fields are evidence based.

Sure, many are not.

But to make such a blanket statement is absurd.


What you have to remember is that the people who decide who to appoint into these roles are another step removed from the problem domain. They will have no programming experience, or the ability to evaluate programming experience. They will try and rely on metrics or some "data-driven" approach to deciding who to appoint.

What kind of data will they look towards? Here, we have a precedent from frantically points to absolutely everywhere around us. So the people with "engagement" and "reputation", i.e. they gushed on their blog and farmed engagement, will be exactly who gets appointed to make the decisions.


The important part of the professional certification like PE or MD is the liability that comes with it, which means that the professional has a lot more agency. A construction firm or hospital can't really force their professionals to do anything because the consequence is criminal prosecution (and a lot of legal liability for the employer).

Bring a software engineer into a courtroom as an expert witness, and the jury's eyes will glaze over. Bring in the PE who told their firm not to cut that corner, and the hammer comes down hard.

Even if the certification for software engineers starts as barebones as knowing what WASP is, it still provides an avenue for the feedback mechanism to work (the rules "written in blood"), so that the entire industry can study and learn from what happened, instead of this mess we have now, the peak of which is postmortem blog posts. Even now we have plenty of examples of regulatory frameworks where the regulations adapt to the field like the FDA where you've got a huge spectrum ranging from diagnostics to medical devices of which where are many classes, and drugs where every clinical trial can be tailored to the exact nature of the disease.


Wasp the defunct PL group from washington.edu or wasp as in owasp or something else entirely?


I have been a professional software developer for decades and worked in several roles where quality and security were at a premium but I have no idea what you mean by WASP. And yet you described it as "barebones" - an interesting illustration of the problem here. Real engineers have many years seeing the hard way what actually works. In software we don't have that kind of consistency and shared understanding of how to reliably get good results yet.


In my experience, good software engineers are often better "engineers" than those working in more physical disciplines, because they get a lot more practice. But traditional engineers have a far better culture when it comes to testing and validation. Since the cycle time is longer, and the cost of mistakes is higher, analysis and testing are generally baked in from the beginning. In software, it's easier to skip that stuff. But software engineers who do get indoctrinated into that culture learn most of the same lessons that traditional engineers do (i.e. how can a component be tested and maintained? what makes for a good design?) and the speed of development means they get more exposure to those types of challenges in general.


In my experience this works both ways too. Good engineers working on physical projects have picked up on some of the useful practices that good software developers have adopted for maintaining progress in the face of ambiguous requirements until they can be clarified and allowing as much flexibility as possible without necessarily compromising quality in the meantime. I believe these relationships can be summarised as something like "Good people keep open minds and learn from what has worked well for others".

The problem for regulating software development is still who gets to formally determine who the "good" people are. This kind of thing should clearly be objective and evidence-based but what useful evidence do we have available?

In physical engineering disciplines there are often clearly evident problems if something was built without being adequately specified by the responsible engineers. In a disastrous case a bridge might literally fall down but you're also going to see that a bridge wasn't designed properly if it's distorting in ways it shouldn't under loads that it should be able to support. There are lots of experienced engineers who have proven records specifying buildings or planes or ships that need to not break using established and peer reviewed techniques.

In software we can all agree catastrophic failures that result in loss of life or half the Internet going down are obviously bad. For something controlling a life-saving medical device or the launch authorisation system for the nuclear missiles we can probably all agree that the answer to what quality level we want in the software is "the best quality we can achieve". But those systems have unusually serious consequences if anything ever goes wrong and probably also very high development budgets that can justify such an extreme position on quality. In general we don't have clearly defined levels of software where different trade-offs between cost and risks and other factors might be considered reasonable and acceptable. Nor do we have well tested and universally accepted standards for how to reliably achieve a specified quality level.


> The problem for regulating software development is still who gets to formally determine who the "good" people are. This kind of thing should clearly be objective and evidence-based but what useful evidence do we have available?

There is zero useful evidence because there is no one to collect it.

The Institution of Civil Engineers was founded in 1818, after decades of random civil engineering societies in Britain doing the exact same thing we are now (running around like chickens with their heads cut off). It wasn't until after the ICE's Royal Charter a decade later that civil engineering began to get really systematized into the "real engineering" we know today and that charter effectively established them as a regulatory body that allowed that to happen.


There is zero useful evidence because there is no one to collect it.

I'm not sure that is entirely true. There have certainly been a few people who have attempted to study what did or didn't work in industrial settings - either pure academics or people working in industrial research labs. But I agree that currently we have nowhere near enough data to form robust conclusions about almost anything in this field and I think this is the strongest argument that the industry is not ready for any kind of licensing and regulation regime.


Once I was talking about a video game and mentioned a global variable holding a reference to the player character and got the eyes wide open horrified look.

Those people would be writing the regulations.


I am a FP nerd and I approve this message


Sadly AI also amplifies both the good and the bad managers and tech leads. If you have managers who fail to set a clear direction or tech leads who fail to define a good framework for developers to operate within then AI just means the poorly directed developer effort goes further in the wrong direction faster. Meanwhile teams with well-specified goals and better structure and processes can exploit the benefits of AI tools much more effectively.


There are armed police in the ticket hall. What do you do? Only 31 seconds now...

Is the suspect black?

(I'm asking this because there is evidence suggesting current facial recognition technology is less accurate at matching faces from some demographics, including black people. Why, what did you think I meant?)


Is this still an issue though? I'm not saying it isn't an issue, but this has been widely cited for years now and I find it hard to believe something hasn't been done about that.


LLMs are inherently dangerous tools

I don't see how. An LLM just generates a stream of output and they became very useful doing no more than that.

What is dangerous is then interpreting that output as instructions to some other part of a system that has the ability to do damage if misused.

and reviewing individual commands (or spamming `y`) doesn't make them less so.

Surely if you review each instruction in the output and do not allow the other part of the system to act on one if it would be harmful then this arrangement is very much less dangerous?


I've caught Fable discovering the ip to a production server in documentation and attempting to connect there on its own to run commands without explicitly being prompted to. It didn't work because I was watching it live and and also the key was password protected, but yeah, I do see some danger.


I have noticed that Fable tends to macgyver solutions together to achieve some goal.


Not only fable. Opus does this too. Which is exactly why I want to review. Like recently for some task it was convinced in a site dump images are not there and convinced itself db and files were skewed. But it didn’t check the actual site … if I hadn’t stopped it, it would have fine on and on or wasted tokens on some elaborate ‘fix’.


My point is that an LLM can't attempt to connect to anything by itself. All an LLM does is produce a stream of output tokens - and that was already quite useful as a coding aid.

It is the harnesses that some people are now wrapping around LLMs to interpret the output from a model as commands to run (or other executable instructions) that are creating all these new risks. Remember that this is still a very recent development and still more recently amplified by the use of feedback loops and long-running agents intended to operate with minimal human supervision.

It is going to be increasingly important to understand exactly what these tools are doing and why for both correctness and security reasons. Not conflating their capabilities with the underlying model that purely generates data is pretty fundamental here.


Once you're running a model inside the harness... you've got yourself a controller inside a control loop, which is genuinely a different kind of thing than just the model alone.

Are you objecting to terminology here?

Are you proposing we say "Fable-In-Claude-Code tried..." instead?

Hmmm... something like that might be necessary. Sure we should typically be tolerant of loose language; but people do keep referring to wildly different contexts in ai conversations, and end up talking past each other.

Running gemini on web is a genuinely different experience to running Fable in claude code, different again from GPT-5.6 in openclaw, or in an ide or etc ...


Yes - I'm objecting to the lazy use of terminology here. LLMs are useful in their own right and are not the real problem here. The real problem is people placing too much trust in inherently unreliable output and then trying to automate away their responsibility to check that output properly before using it.


Depending on the company, that sounds like a bad environment more than a agent issue, no dev/prod network isolation?


> do not allow the other part of the system to act on one if it would be harmful

Network security is really easy right, just don't act on harmful requests


Yeah, just drop when you see the RFC3514 evil bit


If you don't understand clearly what an action proposed by your tool is going to do then why would you permit it?


The article is about measurements taken on this.

One important reason is due to Permission Fatigue: Of course you check everything! You're diligent! The last 100 requests were all ok, so you're down to hitting yes, yes, yes, yes, yes, yes, yes, yes ...

... oops, that third yes should have been a no!


This seems like a problem with the level of abstraction the user interface is working at. It is highly unlikely that in any real world task lasting less than one day there were really hundreds of distinct decisions that needed to be made by the user about appropriate actions to be taken by the agent/harness. It is also highly unlikely that the problem of decision fatigue seen here is somehow magically different to the same problem that countless UI designers had encountered and designed around in other systems long before harnesses running LLMs came along.

This is unfortunately the kind of result you get when you eliminate skilled and experienced people with real understanding of their field and replace them with repeated automatically-generated attempts to solve the same problem until something meeting some basic standard of correctness is found. It's as if the story of agentic AI as it exists today had been compressed into one perfect example of what it can do that is good but also why it's still fundamentally flawed.


I think we agree that there's an interface problem. But there's an intelligence-complete problem hiding underneath; which is why the first instinct was to recruit the human-in-the-loop in the first place. Turns out the human has one of those unintuitive failure modes that occurs when the system gets past a certain level of reliability.

Meanwhile, let's leave the hobby-horses in the closet for now. I won't comment on people's programming tool preferences.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: