Great post, and cudos to both of you, Joel and Karen. Very thorough and well explained. I agree with your partnership model. You are responsible for your own content. How you get there, I am not really worried about. It will be interesting to see how the "ethics" of all this will shake out. Thank you for taking the time to do this and post it.
Thanks so much! Agree this is an ethical gray area. Ultimately, I’d love to see more transparency (and less stigma) around writing with some level of AI assistance.
It’s also important to remember that people (via content farms) have generated some pretty awful content all on our own. 😆
Nice post, you two! Your finding that simple, iterative prompting with a writing sample often outperforms expensive tools is both counterintuitive and practical. It suggests the real skill isn't finding a magic bullet, but in crafting a better, more collaborative process with the AI.
Thank you! ❤️ I was definitely surprised when results from the longer, more specific prompts weren't significantly better. Of course, with non-deterministic models, YMMV.
Great detailed article Joel and Karen again sure will tap into the frustrations of many. The breakdown is seriously impressive -proper rigour and clarity in a space that’s often clouded with jargon and hype. I haven’t tried your humaniser yet as I don’t use Claude!
Thank you! 🤗 For me, one of the biggest takeaways was that "humanized" writing that could bypass AI detection was not necessarily good, publishable writing.
The "reduce editing time, don't eliminate oversight" framing is the right one, and I think it hides a cost the test couldn't see. Your sample was a 250-word blog intro about leash training. Nothing in it renumbers.
Run the same four prompts over a chapter with numbered references, footnotes and a table of contents and the editing you're measuring stops being line edits. It turns into rebuilding. Getting a file into a chat window usually means flattening it first, and citation fields, cross-references and table cells don't reliably come back the way they went in. That's my experience anyway. I've never found a prompt that fixes it, mostly because it isn't a prompting problem.
If you ever revisit this: same four prompts, but score the rebuilt Word file rather than the text you pasted.
I'm on the HumanPen team, we take DOCX and PPTX in and hand the same file back, so that rebuild step is the one I keep looking at.
Great post, and cudos to both of you, Joel and Karen. Very thorough and well explained. I agree with your partnership model. You are responsible for your own content. How you get there, I am not really worried about. It will be interesting to see how the "ethics" of all this will shake out. Thank you for taking the time to do this and post it.
Thanks so much! Agree this is an ethical gray area. Ultimately, I’d love to see more transparency (and less stigma) around writing with some level of AI assistance.
It’s also important to remember that people (via content farms) have generated some pretty awful content all on our own. 😆
Nice post, you two! Your finding that simple, iterative prompting with a writing sample often outperforms expensive tools is both counterintuitive and practical. It suggests the real skill isn't finding a magic bullet, but in crafting a better, more collaborative process with the AI.
💯 I was honestly surprised that the long prompts weren’t obviously better!
Thank you!
Very rigorous examination of this pervasive phenomenon. Interesting insight that more detailed prompts doesn’t necessarily lead to better results.
Well done, Joel and Karen!
Thank you! ❤️ I was definitely surprised when results from the longer, more specific prompts weren't significantly better. Of course, with non-deterministic models, YMMV.
I’ve had the same impression but haven’t done this kind of testing to back it up.
As “non-deterministic” as LLMs may be, they’re very determined to sprinkle em-dashes everywhere they can 😂
Great detailed article Joel and Karen again sure will tap into the frustrations of many. The breakdown is seriously impressive -proper rigour and clarity in a space that’s often clouded with jargon and hype. I haven’t tried your humaniser yet as I don’t use Claude!
I think it’s time to get a free Claude account 🤓
On it 😂
🥳
Thank you! 🤗 For me, one of the biggest takeaways was that "humanized" writing that could bypass AI detection was not necessarily good, publishable writing.
Why do you optimize your text for AI detection, instead of focusing on people?
It looks like that you use AI to tweak AI generated text, so another AI detector, run on AI model will classify text as "Non-AI".
The "reduce editing time, don't eliminate oversight" framing is the right one, and I think it hides a cost the test couldn't see. Your sample was a 250-word blog intro about leash training. Nothing in it renumbers.
Run the same four prompts over a chapter with numbered references, footnotes and a table of contents and the editing you're measuring stops being line edits. It turns into rebuilding. Getting a file into a chat window usually means flattening it first, and citation fields, cross-references and table cells don't reliably come back the way they went in. That's my experience anyway. I've never found a prompt that fixes it, mostly because it isn't a prompting problem.
If you ever revisit this: same four prompts, but score the rebuilt Word file rather than the text you pasted.
I'm on the HumanPen team, we take DOCX and PPTX in and hand the same file back, so that rebuild step is the one I keep looking at.
That would be an interesting expanded test
Your work rate is insane! I don't know how you do it Joel ✨ inspiring. I'm v thankful as I get to enjoy your work 😁
Much of this was Karen, but thanks! 😊