Benchmark Testing AI for Meditation
Domenic Ashburn · Aug 8, 2026
Selecting a generation model for guided meditation from primary contemplative sources. A Sacred Act research report, version 1.0. Study conducted 7 August 2026.
Abstract
We tested eleven AI model configurations from three providers to answer one question: which of them writes the best guided meditation from an ancient technique? The task is exacting. A model must teach a specific practice, drawn from a named primary source with its own mechanism and sequence, to a specific person who has said what they are carrying. No public benchmark measures this. We built one.
Three techniques were selected from published primary sources separated by tradition, geography, and roughly eight centuries: mettā bhāvanā as set out in the Theravāda Vimuttimagga (1st–2nd c. CE), the discipline of testing impressions from the Discourses of Epictetus (c. 108–135 CE), and a fixation practice from verse 113 of the Kashmir Śaiva Vijñāna Bhairava Tantra (c. 8th–9th c. CE). Each was paired with a practitioner scenario for which that practice is indicated. Eighty-one generations were scored blind on five dimensions by two judges from different model families, alongside deterministic measurements of length and format compliance.
We report three results. First, gpt-5.6-terra at low reasoning effort ranked first (8.67/10), and increasing reasoning effort produced no reliable improvement in any provider — in the winning model it produced a decline. Second, an intervention study found that instructional prose in our own prompt was being reproduced verbatim in 88% of generations and was measurably suppressing fidelity to the source technique; removing it improved every model tested, including the incumbent. Third, a defect that rendered generated audio unusable was invisible to all five scored dimensions, which we report as a limitation of the method rather than a finding of the study.
A note on what is disclosed. Methods, measurements, and findings are reported in full, as are the primary sources under test and every model configuration and price. Implementation specifics are not: prompt phrases are identified by role rather than quoted, the structure of the generation prompt and the pacing system beneath it are described functionally rather than specified, and the corpus and practitioner model those systems draw on are outside the scope of this report. Real practitioner scenarios are reported only in aggregate; their content is withheld. Everything needed to evaluate the reasoning is here. Everything that is ours to keep is not.
1. Introduction
The Sacred Act generates guided meditations from a corpus of primary contemplative texts. A generation request supplies a specific technique — not a genre label like "breathing meditation" but an authored practice with its own mechanism, sequence, obstacles, and completion markers, drawn from a named source — together with what the practitioner has said about their present condition. The system must produce a script that teaches that practice to that person.
1.1 Why this study matters to us
The Sacred Act exists to scale awakened awareness by making the practices of humanity's contemplative lineages accessible and personal to the modern mind. Those practices reached us because people across many centuries treated their transmission as a duty: copying, teaching, translating, and refusing to let a method blur into an approximation of itself. When we render one of those practices through a language model, we enter that line of transmission and inherit its obligation. What a practitioner receives must be the practice the source actually teaches, not a plausible imitation assembled from a model's general familiarity with meditation.
This is the responsibility of stewardship in the age of technology, and it sets the terms of the study. The question is not which model writes the most pleasant meditation. It is which model can carry a specific method, intact, from a text written eighteen centuries ago to a particular person sitting down tonight — and can do so while meeting that person where they actually are, because a practice delivered to no one in particular is not personal, and a practice that abandons its source to be personal is no longer that practice.
Both halves of that requirement had to be measured, and measured separately, which is why source fidelity and practitioner attunement are scored as independent dimensions in Section 2.4. A system that quietly substitutes generic mindfulness for an authored method has failed at the thing it exists to do, however well its output reads. Stewardship is either a measurable property of the system or it is only a claim we make about it.
1.2 Why existing benchmarks do not answer this
This task has no public benchmark, and general capability benchmarks do not predict performance on it. A model at the frontier of mathematical reasoning may still produce a script that names a technique and then teaches generic mindfulness under its banner. The failure is not one of intelligence but of instruction-following under a particular kind of constraint: the source is authoritative, the practitioner is specific, and the model's own fluency is the principal threat to both.
We had operated a single model since the product's first version, selected when the field offered few alternatives. Two further provider families had since become competitive. The question requiring an answer was narrow and practical — should we change models, and to which — and could not be answered from published evaluations.
1.3 Research questions
- Which model configuration produces the highest-quality meditation from a primary source technique, at a cost and latency viable in production?
- Does increased reasoning effort improve output on this task?
- Does the ranking obtained from constructed test scenarios hold on real practitioner data?
A fourth question arose during the study and is reported in Section 4: how much of the observed variance is attributable to the models at all, rather than to the prompt they were given.
2. Method
2.1 Task definition
Each configuration received the byte-identical prompt assembled by our production generation route. No prompt was tuned to any model's preferences. The prompt supplies the authored technique record, the practitioner's own words, their prior context, and the environmental conditions of the sitting. The expected output is a complete spoken script with structural pause annotations, targeted at a stated duration.
2.2 Source material
Three techniques were selected to be maximally distant from one another in tradition, method, and idiom, so that fidelity to the source could not be satisfied by any single register of contemplative language.
Table 1. Primary sources under test.
| Technique | Primary source | Date | Tradition | Method type |
|---|---|---|---|---|
| Mettā bhāvanā (graduated) | Vimuttimagga (The Path of Freedom), Ch. VIII | 1st–2nd c. CE | Theravāda Buddhist | Affective cultivation |
| Testing the impression (phantasia) | Discourses II.18; Enchiridion 1, 20, 45 | c. 108–135 CE | Greek Stoic | Cognitive examination |
| Fixed unblinking gaze (stabdha-dṛṣṭi) | Vijñāna Bhairava Tantra, v. 113 | c. 8th–9th c. CE | Kashmir Śaiva (Trika) | Sensory fixation |
These differ in what a faithful rendering even is. The Vimuttimagga prescribes a graduated sequence of recipients and enumerates specific antidotes to hatred; a faithful mettā script must follow that order and reach those antidotes. Epictetus prescribes an examination — interrogate the impression, determine whether it concerns what is in your power, withhold assent from what is not — and a faithful Stoic script must actually perform that operation rather than describe calm. The Vijñāna Bhairava prescribes an unblinking gaze held until the visual field alters; a faithful script must give the stipulation of not blinking, which is the mechanism itself.
2.3 Practitioner scenarios
Each technique was paired with a constructed practitioner for whom that practice is genuinely indicated: a person carrying resentment after a family rupture (mettā), a person in anticipatory catastrophizing three hours before a high-stakes meeting (Stoic examination), and a person whose attention has been fragmented by prolonged screen use (fixation). Each scenario carried stated emotions, the practitioner's own words, a temperament and practice-maturity profile, a duration, and time-of-day context.
Pairing a distant technique with a specific condition allows source fidelity and practitioner attunement to vary independently. A model can teach the Vimuttimagga's sequence correctly while ignoring the person, or attend closely to the person while substituting generic loving-kindness for the authored method. The rubric is designed to separate these.
2.4 Evaluation rubric
Table 2. Scored dimensions. Each 0–10; 5 denotes adequate, 7 good, 9+ exceptional.
| Dimension | Criterion |
|---|---|
| Source fidelity | Teaches this technique's actual mechanism, sequence, vocabulary, obstacles, and completion markers. Penalized for substituting generic mindfulness, omitting core steps, or inventing steps absent from the source. |
| State attunement | Meets this practitioner's stated words, emotions, temperament, growth edge, and hour. Penalized for boilerplate acknowledgment or for referencing what they did not say. |
| Temperament | Warm, unhurried, precise; honors the tradition without pastiche; no coach-register, no over-promised transformation. |
| Structure and pacing | Silence placed where practice occurs; coherent arc; pacing appropriate to duration. |
| Safety | No contraindicated instruction, mishandled difficulty, or medical claim. |
2.5 Judging protocol
Each script was scored independently by two judges from different model families — Gemini 3.1 Pro and Claude Opus 5 — blind to which model produced it. Both judges received the full authored technique record, so that source fidelity was assessed against the primary source rather than against the judge's own knowledge of the tradition. Judges were instructed to score strictly and were given the anchor definitions in Table 2.
Two families were used deliberately. A single judge cannot distinguish a model's quality from its own family's stylistic affinity, and we treat neither judge as authoritative.
2.6 Deterministic measures
Independently of the judges, every script was measured for word count against its pacing target (expressed as a length ratio, where 1.0 is on target), for compliance with the structural annotation rules, and for formatting artifacts that must never reach a speech synthesizer. These measures require no adjudication.
2.7 Models under test
Eleven configurations across three providers. Configurations whose cost or latency would be unviable for a single production generation — frontier-tier models at maximum reasoning effort — were excluded by design. This is an evaluation of production candidates, not a study of capability ceilings.
2.8 Reporting conventions
Scores are means across the three scenarios and both judges. Latency is wall-clock time to a complete script, measured with one request in flight per provider so that no configuration is penalized by our own concurrency. Costs are provider list prices as of 7 August 2026. Absolute scores are not comparable to other benchmarks; only ordering within this rubric is meaningful.
3. Experiment 1: Model selection
Eleven configurations × three scenarios × two judges. All 33 generations completed without failure.
Table 3. Primary results, ranked by quality.
| Configuration | Quality | Latency | Cost/gen | Length ratio |
|---|---|---|---|---|
| gpt-5.6-terra (low) | 8.67 | 12.4 s | $0.0217 | 1.39 |
| claude-opus-5 (low) | 8.63 | 16.1 s | $0.0697 | 1.02 |
| gpt-5.6-terra (medium) | 8.47 | 11.7 s | $0.0215 | 1.30 |
| claude-sonnet-5 (high) | 8.43 | 16.1 s | $0.0416 | 1.12 |
| claude-sonnet-5 (low) | 8.27 | 15.5 s | $0.0406 | 1.06 |
| gpt-5.6-luna (low) | 8.23 | 7.2 s | $0.0020 | 1.36 |
| gemini-3.6-flash | 7.53 | 16.7 s | $0.0330 | 1.07 |
| gemini-3-flash (high thinking) | 7.50 | 10.9 s | $0.0089 | 1.38 |
| gemini-3-flash (minimal) — incumbent | 7.37 | 5.8 s | $0.0055 | 1.51 |
| claude-haiku-4.5 | 6.80 | 12.9 s | $0.0111 | 1.67 |
| gemini-3.5-flash | 6.67 | 15.3 s | $0.0412 | 0.95 |
3.1 Per-dimension results
Table 4. Dimension scores for the six highest-ranked configurations and the incumbent.
| Configuration | Source fidelity | Attunement | Temperament | Structure | Safety |
|---|---|---|---|---|---|
| gpt-5.6-terra (low) | 8.50 | 8.67 | 8.33 | 8.33 | 9.50 |
| claude-opus-5 (low) | 9.00 | 8.33 | 8.17 | 8.17 | 9.50 |
| gpt-5.6-terra (medium) | 8.67 | 8.33 | 8.00 | 8.00 | 9.33 |
| claude-sonnet-5 (high) | 8.67 | 8.33 | 8.00 | 8.17 | 9.00 |
| claude-sonnet-5 (low) | 8.17 | 8.33 | 7.83 | 7.67 | 9.33 |
| gpt-5.6-luna (low) | 8.00 | 8.17 | 7.67 | 8.00 | 9.33 |
| gemini-3-flash (incumbent) | 8.00 | 6.67 | 6.83 | 7.00 | 8.33 |
The incumbent's deficit was concentrated: its source fidelity (8.00) was competitive with mid-ranked configurations, but its attunement (6.67) was the lowest score in the top nine. Qualitatively, its scripts summarized the practitioner's stated condition rather than engaging it, and reached for elevated register where the leaders used the practitioner's own vocabulary.
Judge commentary on the failures was specific to the sources. On the Vijñāna Bhairava scenario, one judge noted that a script "never actually instructs the unblinking quality that is the verse's core stipulation" — a script that reads as a competent gaze meditation while omitting the mechanism the verse exists to transmit. On the Vimuttimagga scenario, judges repeatedly penalized compression of the graduated sequence and omission of the text's enumerated antidotes to hatred. These are the failures the rubric was built to detect, and they are invisible to any general benchmark.
3.2 The effect of reasoning effort
Within each provider we tested at least two reasoning settings.
Table 5. Quality as a function of reasoning effort.
| Model | Lower setting | Higher setting | Δ |
|---|---|---|---|
| gpt-5.6-terra | 8.67 (low) | 8.47 (medium) | −0.20 |
| gemini-3-flash | 7.37 (minimal) | 7.50 (high) | +0.13 |
| claude-sonnet-5 | 8.27 (low) | 8.43 (high) | +0.16 |
No provider showed a reliable gain. Both positive deltas fall within the sampling noise of a three-scenario design; the single negative delta occurred in the highest-scoring model. Every increase carried a real cost in latency and tokens.
We interpret this as a property of the task rather than of the models. Composing a meditation from an authored technique is not a problem to be solved prior to being written: the source supplies the sequence, the practitioner supplies the occasion, and what remains is execution in a voice. Extended deliberation appears to contribute nothing that a model's first fluent register does not already provide, and may plausibly cost something if deliberation biases the output toward explanation over instruction.
All production configurations were subsequently pinned to their lowest reasoning setting.
3.3 Cost and quality
Table 6. Cost against quality at the extremes.
| Cost/gen | Quality | |
|---|---|---|
| Most expensive tested | $0.0697 | 8.63 |
| Highest scoring | $0.0217 | 8.67 |
| Least expensive tested | $0.0020 | 8.23 |
| Lowest scoring | $0.0412 | 6.67 |
Cost was a poor predictor of quality within this task. The most expensive configuration did not rank first; the least expensive outranked a configuration costing twenty times more. Model fit to the task dominated model tier.
3.4 Inter-rater reliability
Table 7. Judge calibration across all eleven configurations.
| Value | |
|---|---|
| Mean score, Judge A (Gemini 3.1 Pro) | 8.48 |
| Mean score, Judge B (Claude Opus 5) | 7.26 |
| Mean offset | 1.22 |
| Rank correlation (Spearman ρ) | 0.81 |
The judges differed substantially in absolute calibration and agreed strongly on ordering. Both placed gpt-5.6-terra (low) and claude-opus-5 (low) in their top three. Neither judge belonged to the provider family of the winning configuration, so the result carries no same-family advantage.
This is the principal methodological caveat of the study. A single-judge design would have shifted every reported score by more than a full point depending on which family was selected, while leaving the decision unchanged. We report rankings as robust and absolute scores as internal to this rubric.
4. Experiment 2: The prompt as a confound
Inspection of the generated scripts — rather than of their scores — surfaced a systematic effect that the ranking had absorbed rather than isolated. Instructional prose in our own prompt was being reproduced in the output.
The clearest instance was a paragraph of posture guidance written years before our source corpus existed, when the system was required to supply more of the practice itself. A single distinctive phrase from that paragraph appeared in 29 of 33 first-round generations (88%). Five further sites in the prompt produced the same effect: illustrative sentences offered as examples of tone, atmospheric description intended as guidance, and abstract section headings whose vocabulary surfaced in roughly a quarter of all scripts.
The consequence was not stylistic. The same prompt instructed the model to follow the source technique's own preparatory instructions, and then supplied a competing set — a generic seated-posture register applied identically to the Vimuttimagga's affective cultivation, Epictetus's cognitive examination, and the Vijñāna Bhairava's sensory fixation. Three practices that agree on nothing were being set up identically.
4.1 Intervention
Every such site was rewritten to state intent and constraint in place of language. Where the prompt had supplied a sentence, it now supplies a goal and its boundaries and leaves the wording to the model, the source, and the practitioner. Illustrative examples of desired phrasing were removed; examples of phrasing to avoid were retained, as those are not copied. An explicit precedence order was added, so that where instructions conflict, the source technique and the practitioner's present words govern.
The models, prompt-independent factors, judges, scenarios, and rubric were held constant. Rounds 2 and 3 re-ran four configurations — the three highest-ranked plus the incumbent — before and after the rewrite.
4.2 Results
Table 8. Reproduction of prompt phrases in generated output, by round. Phrases are identified by role rather than quoted.
| Prompt phrase | Round 1 | Round 2 | Round 3 (post-rewrite) |
|---|---|---|---|
| A — distinctive anatomical term | 29/33 | 19/24 | 0/24 |
| B — posture adjective | 19/33 | 10/24 | 2/24 |
| C — instruction fragment | 16/33 | 11/24 | 0/24 |
| D — atmospheric phrase | 15/33 | 5/24 | 0/24 |
| E — anatomical fragment | 9/33 | 7/24 | 0/24 |
| F — abstract noun from a section heading | 8/33 | 7/24 | 0/24 |
| G — unusual adjective from an example sentence | 3/33 | 2/24 | 0/24 |
Table 9. Quality before and after the prompt intervention.
| Configuration | Pre-intervention | Post-intervention | Δ |
|---|---|---|---|
| gpt-5.6-terra (low) | 8.47 | 8.80 | +0.33 |
| claude-sonnet-5 (low) | 8.40 | 8.50 | +0.10 |
| gpt-5.6-luna (low) | 8.43 | 8.47 | +0.04 |
| gemini-3-flash (incumbent) | 7.83 | 8.00 | +0.17 |
Reproduction fell to zero for six of seven tracked phrases, and quality improved for all four configurations. Length ratios converged toward target across the board; the incumbent moved from 1.51 to 1.33.
The result is a confound rather than an incidental finding. Under the original prompt, a portion of what the ranking in Table 3 attributed to model quality was in fact each model's differing propensity to reproduce supplied text. The ordering of the top configurations was unchanged by the intervention, so the selection conclusion stands; the magnitude of the gap between models was overstated.
4.3 Regression control
Prose discipline is not self-sustaining. Every generated script is now compared against its own prompt with all technique- and practitioner-specific content removed, and any shared four-word sequence is flagged. One phrase is exempt by design, being a required element of the product's voice. Round 3 produced zero flags across 24 scripts, and any future prompt revision that reintroduces supplied language fails this check automatically.
5. A limitation of the method
Both judges scored one script 8.5 on structure and pacing. Rendered to speech, its central passage failed: a sequence of short lines intended to be received individually was delivered as a single unbroken run, faster than any of them could be received.
The cause lay in the layer between the script and the synthesizer, and was invisible in the text. The script was correct. The rendering was not. No dimension in Table 2, applied to text, could have detected it, and it was found only because the audio was produced and listened to.
We report this as a limitation of the study design rather than as a result. A rubric measures what it measures. Where the delivered artifact is heard rather than read, an unbroken record of high text scores is weaker evidence than it appears, and rendering into the final medium is not an optional final step but part of the measurement.
6. Experiment 3: Validation on practitioner data
Constructed scenarios may flatter a model in ways that real requests do not. Rounds 2 and 3 added three sittings drawn from a live practitioner account, carrying the full context a production request carries. Scenario content is withheld; only aggregate scores are reported.
Table 10. Constructed versus real practitioner data, post-intervention.
| Configuration | Constructed | Real account | Δ |
|---|---|---|---|
| gpt-5.6-terra (low) | 8.80 | 8.57 | −0.23 |
| gpt-5.6-luna (low) | 8.47 | 8.37 | −0.10 |
| gemini-3-flash (incumbent) | 8.00 | 7.97 | −0.03 |
| claude-sonnet-5 (low) | 8.50 | 8.00 | −0.50 |
The leading configuration held its rank. The decisive observation came one round earlier, under the pre-intervention prompt: gpt-5.6-luna, the least expensive configuration in the study at approximately one-tenth the leader's cost, scored 8.43 on constructed scenarios — 0.04 behind the leader — and then fell to 7.83 on real data, the largest degradation of any configuration, with source fidelity declining most sharply.
On constructed evidence alone, a near-tie would have justified selecting a configuration costing an order of magnitude less. Real data reversed that conclusion. We regard this as the strongest argument in the study for validating on production inputs before deciding: constructed scenarios are cleaner than real requests in precisely the respects that distinguish models.
7. Decision and deployment
gpt-5.6-terra at low reasoning effort was deployed as the primary generation engine on 7 August 2026, with the previous engine retained as an automatic fallback: an absent key, an open circuit breaker, or a failed call all revert to the prior path unchanged.
Table 11. Deployed configuration against the incumbent, measured on real practitioner data.
| Incumbent | Deployed | Δ | |
|---|---|---|---|
| Quality | 7.97 | 8.57 | +0.60 |
| Source fidelity | 7.50 | 8.50 | +1.00 |
| State attunement | 7.50 | 8.50 | +1.00 |
| Temperament | 7.83 | 8.50 | +0.67 |
| Latency | 4.9 s | 13.2 s | +8.3 s |
| Cost per generation | $0.0058 | $0.0238 | +$0.0180 |
The additional cost is a small fraction of the total cost of producing one meditation, which is dominated by speech synthesis rather than text generation. The additional latency occurs within a pipeline in which synthesis dominates wall-clock time.
8. Threats to validity
Sample size. Three scenarios per round. Differences below approximately 0.3 points should not be treated as meaningful. The reasoning-effort deltas in Table 5 fall within that band; that conclusion rests on the consistent absence of gain across three providers rather than on any individual measurement.
Judges are language models. They share failure modes with the systems under evaluation and may reward fluent prose over prose that would serve a practitioner in difficulty. The two-family design and the reported 1.22-point calibration gap are mitigations, not solutions.
Scores are internal. Absolute values are meaningful only relative to one another within this rubric.
Source selection. Three techniques cannot represent a corpus spanning many traditions. They were chosen for maximal separation, which tests breadth but may understate performance on practices closer to a model's training distribution.
Prices are dated. List prices as of 7 August 2026 and subject to change.
No outcome data. We measure the quality of a script, not whether any practitioner settled. That is the more important question and this study does not address it.
Method blindness. As reported in Section 5, the defect with the largest effect on a listener was undetectable by the method described in Sections 2 through 4.
9. Conclusions
gpt-5.6-terraat low reasoning effort produced the highest-quality meditations from primary contemplative sources among eleven production-viable configurations, and its ranking was stable across a prompt revision and a change from constructed to real practitioner data.- Increased reasoning effort did not improve output for any provider on this task. For craft and voice generation, the lowest setting should be evaluated first.
- Instructional prose supplied in a prompt is reproduced in output at high rates, competes with authoritative source material, and depresses measured quality. Its removal improved every configuration tested and represented a larger single-factor gain than the model change itself.
- Evaluations conducted on constructed scenarios can select the wrong model. Validation on production inputs altered the decision here.
- A text-scored rubric can assign high marks to an artifact that fails in its delivered medium.
The study was designed to answer a procurement question and did so. Its more consequential result was the discovery that a substantial share of the variance we had attributed to models originated in instructions we had written ourselves. A benchmark is generally described as an instrument for evaluating models. This one functioned equally as an instrument for evaluating its authors.
81 generations across 11 configurations, 3 providers, 3 primary sources, and 2 blind judges. Zero generation failures. Conducted 7 August 2026.
