I started KeenLore (an emotive audiobook creator) that way. I gave it software specifications, languages, JSON schema definitions, container requirements, hardware configuration (8GB NVIDIA T1000 GPU, 96GB RAM), and zero user interface mockups. For the second round, I asked it to build a completely independent, re-entrant, and data isolated demo system on top of the web application. The demo application included voice generation using one of its voice designs. Here's the output:
https://www.youtube.com/watch?v=WAeHgE94rVo
The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.
The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).