If you're building a Jev-style "System One" model, please add prompt caching or reusable question sets to your API 🙏. This would make batch use cases even more efficient, so you could amortize the cost of many questions asked over the same input. (This was also proposed in typesafe-ai/typesafe-sdk-js#10.)
TL;DR: after a week of tweaking my Jev calls, my questions are ~88% of the input tokens. I'm asking 54 questions of ~1.1 million documents per month. Jev makes certain types of classification tasks easy and cheap, but prompt caching would make batch workflows even more cost effective. For me, the total dollar amount is still reasonable (less than $150 per month), but I'm sure others will hammer these APIs even harder.
I work on Scour. It’s a personalized content feed where you describe topics you’re interested in and it finds articles and blog posts related to them.
For some time, I’ve wanted to ask slightly fuzzy questions of each piece of content Scour pulls in. What level of experience does this assume? Will this still be worth reading after a week, or is it news that goes out of date quicker? However, running the >1 million pieces of content Scour ingests each month through even the cheapest LLM would be prohibitively expensive for a bootstrapped project. System One models make this kind of use case feasible.
Yes. I could. I might. But having a kind of general purpose classifier that I can ask a barrage of questions and add more questions to over time without retraining is very interesting.
I’ve spent a good chunk of the last week running experiments with Jev, tuning questions, comparing Jev's answers to panels of LLM judges and spot checking the results. The questions help me weed out junk that was hard to spot deterministically, figure out the expertise assumed by posts (to hide beginner content from advanced readers), and more.
Following TypeSafe's advice about atomic questions and sending all questions in one request, v28 of my question set includes 54 questions: 50 nouls, 3 choices, and 1 score. I also reworded most of my questions to remove as much of the criteria description as I could without losing too much accuracy on the results.
This set of questions is approximately 2,150 tokens, including Jev’s fixed overhead that seems to be around 260 tokens (judging from the token count for a trivially short question). When I send this set of questions in with the post’s title, URL, and summary or first snippet of text, my questions represent ~88% of my input tokens. The questions are sent verbatim with every request.
Note that I don't send multiple posts in the same request because of the warning that "Accuracy falls as the state grows with content unrelated to the decision."
To TypeSafe and all the other labs that follow with these types of models, please add prompt caching or the ability to register sets of questions and criteria and reuse them.
I don't know enough about how these models are served and whether there are caching efficiencies to be gained on the backend, but this would help increase Jev's effective "intelligence-per-dollar".
With cheaper repeated questions, I'd put back some of the criteria I cut down, and probably add even more questions.
As another alternative, an API that takes multiple questions and multiple inputs and gives you the results for each question evaluated for each input would give some of these savings without taking the accuracy hit from packing multiple inputs into one request.
Async batch mode probably makes sense for super high-volume, fully offline use cases. It's less relevant for me because I want newly ingested content to be available soon after Scour reads it. But others will probably find plenty of use for it.
Keep up the great work! This is a very exciting development.