Selecting fast mode
Three equivalent ways — use whichever your client makes easy:fast: true is accepted as an alias for mode: "fast"; if both are present,
mode wins.
What a fast answer looks like
- Shape: the answer first (a few sentences), then the source — quoted verbatim in the original with a translation and its canonical reference — then at most a one-line takeaway.
- One retrieval pass: fast mode searches the Chabad corpus (Igros Kodesh, Likkutei Sichos, Toras Menachem, Shulchan Aruch HaRav, and more) or resolves a specific reference (a pasuk, a Tanya chapter) directly. It does not run multi-round research.
- Honest limits: if one pass can’t verify a source for your question, the answer says so plainly and suggests a standard-mode turn — it will not pad or invent.
- Depth on demand: for full shakla v’tarya, cross-references, and source
chains, use standard mode (omit
modeor set"standard").
Billing and limits
Fast turns bill actual usage like any other turn —usage.cost on the response
(and the x-ravchat-credits header) reports the exact charge, which is typically
a small fraction of a standard research turn. Rate limits and concurrency follow
your plan, same as standard mode.
Fast turns are best used stateless: send your full conversation in
messages.
Session affinity (metadata.ravchat.session_id) works, but mixing fast and
standard turns in one long-lived session keeps no shared server-side transcript
between the modes.