Overview
We’re introducing Vivral V2 Swift, our fastest variant of Vivral, engineered for edge devices. For the past couple of months Vivral V2 Swift has been tested in our beta tester group and has outperformed Vivral Surreal, our last-generation frontier model, massively in real-world testing and in the benchmarks.
First we show Vivral V2 Swift Base — the base model with no post-training whatsoever, which still punches well above its weight on many medical benchmarks. Then we show production Vivral V2 Swift results on PubMedQA against frontier models.
Vivral V2 Swift Base benchmarks
Below we show the benchmarks we recorded for the Vivral V2 Swift Base model. This model has no post-training whatsoever and punches well above its weight on many medical benchmarks.
Vivral V2 Swift benchmarks
These results are for the production Vivral V2 Swift model — post-trained for real use — not Swift Base above.
At a fraction of frontier model size, Vivral V2 Swift still sits in the same band as leading closed models on clinical accuracy benchmarks. On PubMedQA, Swift scores 74.8%, next to Gemini 3 Pro (75.2%) and ahead of GPT-5 (73.4%), with GPT-5.6 Luna at 78.4%. Edge-scale efficiency without giving up clinical reading ability is what Swift is built for.
Efficiency is part of the product story. Against GPT-5.6 Luna (at a ~1T parameter estimate), production Swift shows far higher medical intelligence per parameter, lower estimated cost per question, and dramatically lower environmental impact per day of use.
How Vivral was trained
Vivral was developed under tightly limited computational resources. The work relied on a small laboratory stack rather than a multi thousand GPU training run, yet still attained near frontier quality on clinical reading tasks that matter for everyday use. The hardware consisted of one NVIDIA DGX Spark, one RTX 3090, one M4 Mac mini with 32 GB of memory, and one M1 Pro MacBook Pro. Across the full program the total compute was on the order of 1021 to 1022 FLOPs, far below the budgets typical of frontier pretraining.
Training proceeded in stages. Pretraining continued for several months and was performed mainly on the DGX Spark. Near the end of that phase, in order to reduce reliance on cloud GPUs, approximately one week of midtraining was run entirely on the Spark. That midtraining stage used on the order of a few billion tokens and emphasized medical domain material. Supervised fine tuning then continued for about one month, again primarily on the Spark.
Reinforcement learning was carried out for roughly two weeks on the DGX Spark together with cloud hosted 3090 GPUs. The objective of this stage was patient facing quality: careful tone, appropriate bedside manner, and answers that are useful in ordinary health conversations. The process was not designed to maximize clinical benchmark leaderboards.
The model was trained for the public rather than as a medical superintelligence for the hardest specialist cases. Priority was given to questions that patients, nurses, medical students, and physicians reported as frequent and important to answer well. Substantial effort went into the kinds of responses people actually need in practice. At its smallest size, Vivral is not intended to solve extreme diagnostic puzzles of the sort popularized by fictional case work. It is intended as a private, affordable layer of Health Intelligence for everyday clarity.
The entire program cost on the order of approximately 5,000 United States dollars. Careful training strategy and architectural choices made competitive results achievable without multi billion dollar data centers or industrial scale training infrastructure. Efficiency was treated as a core operating principle of the company rather than a secondary concern.
Information security
When a user chooses to rate a response, optional feedback may be sent to our servers. That transmission is never automatic. The product always presents an explicit confirmation step on the user device before any rated exchange leaves the device. If the user declines, nothing is uploaded for that rating event.
Before any confirmed rating payload is accepted, the content passes through a personally identifiable information (PII) redaction system that we have tested for high accuracy on names, dates of birth, medical record numbers, contact details, addresses, and related identifiers. Clinical wording that is not identifying is preserved so that the feedback remains useful for product quality work without retaining unnecessary personal data.
The examples below are illustrative only. They show the style of transformation applied by the redaction pipeline when a user has opted in and confirmed the send.
Example 1 · Direct user question
Input
I am Maya Rodriguez. My date of birth is 04/18/1987 and my MRN is MRN-4839201. I have chest pressure and take nitroglycerin. Email maya.rodriguez@example.com or call (415) 555-0182. Should I go to the ER?
Output after PII redaction
I am ⟦NAME REDACTED⟧. My date of birth is ⟦DATE OF BIRTH REDACTED⟧ and my MRN is ⟦MEDICAL RECORD REDACTED⟧. I have chest pressure and take nitroglycerin. Email ⟦EMAIL REDACTED⟧ or call ⟦PHONE REDACTED⟧. Should I go to the ER?
Example 2 · Caregiver question
Input
I am asking for my father, Daniel Kim, who lives at 81 Harbor Street, Boston, MA 02110. His health plan ID is HP-77492061. His blood sugar is 310 despite metformin, and he feels confused tonight. What should we do?
Output after PII redaction
I am asking for my father, ⟦NAME REDACTED⟧, who lives at ⟦ADDRESS REDACTED⟧. His health plan ID is ⟦HEALTH PLAN ID REDACTED⟧. His blood sugar is 310 despite metformin, and he feels confused tonight. What should we do?
In short, rating data reaches our servers only when the user intentionally chooses to share it and confirms that choice. When that happens, the PII system redacts identifying fields before the material is retained for quality review.
Safety and intended use
While Vivral achieves competitive results on medical benchmarks and in real world testing, it is important to state clearly that Vivral is not a medical device and is not a source of diagnosis or treatment. Vivral has not been reviewed or approved by the United States Food and Drug Administration or by any comparable regulatory authority as a medical device, clinical decision support system, or diagnostic product. It must not be used as a substitute for professional medical judgment, emergency care, or a relationship with a licensed clinician.
Vivral is engineered for medical education and health literacy. Its purpose is to help people understand wellness topics and organize questions using a model trained on a large body of educational and clinical style material. The system is designed to provide general orientation and explanation, not to deliver personalized clinical conclusions, prescribe interventions, or manage acute conditions.
Large language models remain prone to hallucinations, omissions, and confident errors. Substantial work has gone into reducing these failures for patient facing use, yet no mitigation eliminates them. Users should treat all outputs as provisional information and should verify any health related guidance with a qualified healthcare provider before acting on it. When symptoms are severe, worsening, or uncertain, seek appropriate clinical care rather than relying on the model.
In summary, Vivral V2 Swift is an educational Health Intelligence product. Competitive benchmark scores and favorable field observations do not constitute clinical validation, regulatory clearance, or a claim of medical device status. Responsible use requires human oversight and professional medical confirmation wherever care decisions are involved.
Availability
Vivral V2 Swift is being tested with our beta group. Request access if you want to join the waitlist for the final release.
