upvote
I'm mostly using deepseek 4.1 flash via openrouter, the way I'm doing stuff is:

1. write a sketch of a spec by hand

2. have the llm review the document and question me until it can generate a spec

3. review the spec and revise where needed

4. have it write an implementation plan

5. another round or revision/review

6. executing the plan step by step through the plan, plausing between each step to see if we are still on course and if the decisions it made track with my understanding of what we are doing.

I've been working for a couple of hours tonight, the total cost of the session is €0.6.

it's not the build this thing end to end, but also not quite write function x for me. It is still a lot of manual review, but I find I really need it to even discover what I actually want to build. I just cannot imagine building something in a single shot and getting something that actually has value (unless it is basically a clone of an existing thing). To me the whole value of ai right now is that it's now very cheap to build custom software that exactly matches your preferences.

reply
One of the big differences using the better models is that you don’t need to hold their hand anywhere near as much. Fable is crazy expensive, but I’ve seen it just one-shot some remarkably complex projects. The code it produces is much better, too.
reply
yeah i just think that if I gave the initial pitch to Fable it would have successfully built something that wasn't exactly what I wanted. maybe it is because i went to school for design, but how something works is so important to me and it is hard to get to that point without actually thinking about it step by step and visualizing how interactions would work. I just cant imagine a single one-shot prompt containing enough information to for the model to know what to build. I guess if you have Frontier model money like these silicon valley freaks you could just iterate by asking it to tweak what you didn't like until it is right.
reply
This was my experience 3 months ago. I had an Android app that interacted with a Bluetooth device that I wanted to reverse engineer and build my own Linux app for it. DeepSeek was struggling really hard. Claude did it end to end after 3 or 4 prompts. To be fair, I was using a web interface for DeepSeek and the CLI for Claude; maybe that makes a large difference.
reply
build this prototype from end-to-end, is this how people build serious software with AI?
reply
You're not gonna get something that's ready to ship, but as a first pass to get something running yes. Let's you explore far more ideas with only a few hours of agent time.
reply
And you can’t use deepseek with a loop approach for creating such PoCs?
reply
I started KeenLore (an emotive audiobook creator) that way. I gave it software specifications, languages, JSON schema definitions, container requirements, hardware configuration (8GB NVIDIA T1000 GPU, 96GB RAM), and zero user interface mockups. For the second round, I asked it to build a completely independent, re-entrant, and data isolated demo system on top of the web application. The demo application included voice generation using one of its voice designs. Here's the output:

https://www.youtube.com/watch?v=WAeHgE94rVo

The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.

The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).

reply
I don't think most prototypes are serious software.
reply
This is how ChatGPT, Cursor apps are built. They spent so much money on PR stunts, but "thousands of agents" can't make an app that doesn't freeze on each keystroke. Not even talking about user-friendly ui
reply
Weird, ChatGPT has always worked really well for me.
reply
Such issues are pretty common among people I talk to at least. I hear comments about it pretty often, both when it comes to ChatGPT and Claude.

For myself, with ChatGPT for example, it regularly gets extremely slow if there is a lot of text in a single conversation. Especially if I try to scroll up.

Some pages have been strangely just broken for a while now as well. Usage analytics just renders lots of these duplicate "Usage history" components, where the data just never loads: https://imgur.com/a/vpcQIiw.png

I can totally imagine a scenario where some agent built it, another tested and approved it, and nobody at OpenAI even looked at it once or knows that it's like this.

reply
Early on their software was like this, the original ChatGPT desktop app was basically unusable and full of memory leaks that would tank the software. They’ve long since fixed that though
reply
The new ChatGPT desktop app is a dumpster fire though, it can’t even scroll properly when streaming the answer.
reply
For a prototype or a 1 or 2 use tool, yes, this is exactly how serious people are building software.
reply
Oh buddy you have no idea what a good plan and agent harness can do with deepseek…
reply
I am doing mostly hardware drivers recently, it works fine for complex work.
reply
"build this prototype from end-to-end" works fine with DeekSeek V4.1 Flash, the problem occurs if you're not only building a prototype but want a finished product.
reply