Replying to an earlier post

It feels so weird to read lemmy, as if I’m living in a different reality. To me, most of the time an agent can oneshot a ticket (if it has a good, non-vague description) or at least do 80-90% that can be fixed with several changes or prompts, and only odd tasks need more manual investigations than that.

Surely you still need a dev oversight and someone needs to do tech plans & lead the projects, but that’s not like anything people experience here.

Replying to an earlier post

Anthropic claims that Fable has low low low hallucination rate of just 40%.

LLMs outputting code solving this exact task is a compound function of luck, with non determined a priori chance of success and unknown a priori cost.

This is my main disappointment with agentic coding. Prompting is fine, quality assurance is there from the start. Greenfield, couldn’t care less. Established enterprise code? This is a minefield.

If the tokens were 10 to 100 times cheaper, then it would be a maybe.

I also hate how it makes half of my senior engineers dumber.

en