SeedVR2 vs Topaz Wonder 3: A 264-Pair Blind Evaluation of Two Image Upscalers
In the research, 13 evaluators compared SeedVR2 and Topaz Wonder 3 on 264 image pairs and tied 54-60% of them. SeedVR2 leads on edge sharpness, Topaz Wonder 3 wins faces, and the free model is the defensible default.
Evaluation period: July 2026 · Session: ff12fe31 · Confidence level: directional
DIRECTIONAL RESULT
Inter-rater agreement in this session (Gwet’s AC1 = 0.06–0.12) falls below the threshold Everypixel requires for a verified label. Every aggregate figure below describes direction, not an attested value. The reason for the low agreement is itself a finding and is reported in Section 3.5.
Abstract
SeedVR2 and Topaz Wonder 3 are statistically close to indistinguishable in blind pairwise evaluation. Across 264 image pairs and 791 scored comparisons by 13 professional evaluators, 54–60% of pairs were scored as ties. SeedVR2 holds a small but interval-separated advantage on edge sharpness (P = 0.60, 95% CI 0.57–0.62) and detail recovery (P = 0.55, CI 0.52–0.57); naturalness is not separated (P = 0.51, CI 0.48–0.54). Overall P(SeedVR2 > Topaz Wonder 3) = 0.55 (CI 0.53–0.58). One category reverses the ranking: on faces, Topaz Wonder 3 wins naturalness on 52% of pairs against 14%. (Everypixel production team, July 2026)
1. Introduction
Two upscalers are compared here: SeedVR2, a one-step diffusion restoration model from ByteDance with published weights that runs locally at no license cost, and Topaz Wonder 3, a paid commercial upscaler with a tunable parameter interface. Both are applied to still-image upscaling of production stock content.
The comparison follows the same protocol Everypixel used in June 2026 to compare SeedVR2 against FLUX.2 [klein] 9B: identical 264-pair structure, identical category split, the same 13-person evaluator panel. That earlier session separated the two models on every scale. This one did not, and the difference between the two outcomes is itself the primary result reported here.
Two consequences follow. First, evaluator agreement in this session fell to Gwet’s AC1 = 0.06–0.12, against 0.42–0.71 in the June session, which places the result below the threshold for a verified label under the internal evaluation protocol. It is published as directional. Second, because the two models are separated by less than the panel can reliably resolve on a single pair, the practical selection criteria shift to cost, category fit, and failure mode rather than aggregate quality.
2. Methods
2.1 Study design
The test set comprised 264 image pairs: 132 real photographs and 132 AI-generated images, distributed across six content categories (crowds and complex scenes, faces, materials and textures, nature, text and signage, urban scenes), 44 pairs per category.
Each pair was presented blind in randomized left/right order and scored in A/B/tie format on three quality scales:
| Scale | Definition |
|---|---|
detail_recovery | Recovery of structure present in the source but not resolved at source resolution |
edge_sharpness | Definition of edges and boundaries without visible sharpening artifacts |
naturalness | Absence of synthetic or over-processed appearance; plausibility of the result as a photograph |
Thirteen evaluators participated across three roles: 4 QA testers, 4 art directors, and 5 distribution reviewers. Each pair was targeted for three independent ratings; 791 of 792 planned comparisons were completed.
2.2 Statistics
Inter-rater agreement is reported as Gwet’s AC1 rather than Fleiss κ. Ties account for roughly 60% of outcomes, and under a distribution this skewed the Fleiss κ expected-agreement term inflates, and κ collapses toward zero independently of actual rater concordance, a known artifact described as the paradox of kappa (Feinstein & Cicchetti, 1990). AC1 is stable under skew (Gwet, 2002). Fleiss κ is reported alongside AC1 for completeness and reads 0.00–0.05 in this session, which is an artifact rather than a measurement.
Aggregate model strength uses the Bradley-Terry model (Bradley & Terry, 1952), fitted per scale and overall, with 95% confidence intervals from bootstrap resampling. Reported probabilities are P(SeedVR2 preferred over Topaz Wonder 3) in a single comparison; a confidence interval containing 0.50 indicates no separation.
Win percentages in Sections 3.1 and 3.2 are pair-level (n = 264 per scale, aggregating the three ratings per pair). Percentages in Section 3.4 are vote-level since role subsets have unequal rating counts (n = 303 distribution reviewers, 244 art directors, and 244 QA testers).
3. Results
3.1 Overall
| Scale | SeedVR2 | Tie | Topaz W3 | AC1 | Fleiss κ | P [95% CI] |
|---|---|---|---|---|---|---|
| Detail recovery | 25% (67) | 60% (159) | 15% (38) | 0.12 | 0.05 | 0.55 [0.52–0.57] |
| Edge sharpness | 35% (93) | 54% (142) | 11% (29) | 0.06 | 0.00 | 0.60 [0.57–0.62] |
| Naturalness | 21% (56) | 59% (156) | 20% (52) | 0.08 | 0.04 | 0.51 [0.48–0.54] |
| Overall | — | — | — | — | — | 0.55 [0.53–0.58] |
Pair-level, n = 264 per scale. P = Bradley-Terry probability that SeedVR2 is preferred in a single comparison. Percentages are allocated by largest remainder, so each row sums to 100; counts are exact, so a cell may sit up to one point off its count. (Everypixel production team, July 2026)
Pair-level, n = 264 per scale. P = Bradley-Terry probability that SeedVR2 is preferred in a single comparison. (Everypixel production team, July 2026)
The overall effect is real but small. A probability of 0.55 corresponds to a preference expressed in roughly one comparison in twenty above chance against a tie rate near 60%.
3.2 By content category
| Category | Detail A/B/tie | Sharpness A/B/tie | Naturalness A/B/tie |
|---|---|---|---|
crowds_complex |
41 / 11 / 48 | 29 / 14 / 57 | 43 / 16 / 41 |
faces |
18 / 25 / 57 | 36 / 21 / 43 | 14 / 52 / 34 |
materials_textures |
25 / 7 / 68 | 68 / 5 / 27 | 16 / 11 / 73 |
nature |
18 / 14 / 68 | 36 / 5 / 59 | 11 / 21 / 68 |
text_signage |
20 / 16 / 64 | 25 / 11 / 64 | 18 / 14 / 68 |
urban_scenes |
29 / 14 / 57 | 16 / 11 / 73 | 25 / 5 / 70 |
A = SeedVR2, B = Topaz Wonder 3. Percentages of pairs within the category (n = 44). Percentages are allocated by largest remainder, so each row sums to 100; counts are exact, so a cell may sit up to one point off its count. (Everypixel production team, July 2026)
A = SeedVR2, B = Topaz Wonder 3. Percentages of pairs within the category (n = 44). (Everypixel production team, July 2026)
Two category effects exceed the noise floor of the overall result.
The face result is consistent with the qualitative reports in Section 4. Topaz Wonder 3 renders facial skin as a continuous, even surface, which raters scored as natural at portrait scale. SeedVR2 resolves pore-level structure and, when the reconstruction is incorrect, places the error on a face. The two models also trade in the opposite direction on a second face attribute: raters reported that SeedVR2 preserves facial proportions and likeness more closely, while Topaz Wonder 3 alters features more visibly. Naturalness and fidelity are not the same measurement, and on faces the two models split them.
3.3 By source type
| Source type | Detail A / tie | Sharpness A / tie | Naturalness A / tie |
|---|---|---|---|
| Real photographs (n = 132) | 29% / 60% | 35% / 56% | 19% / 64% |
| AI-generated (n = 132) | 22% / 61% | 36% / 52% | 23% / 55% |
A = share of pairs on which SeedVR2 was preferred. (Everypixel production team, July 2026)
A = share of pairs on which SeedVR2 was preferred. (Everypixel production team, July 2026)
This is a negative result worth recording. In the June 2026 session against FLUX.2 [klein] 9B, source origin identified the one condition under which the losing model won. Here it carries no decision value, and a mixed-content pipeline does not need to route by source type.
3.4 Between-role variation
Cross-role agreement is correspondingly low: AC1 = 0.18 (detail recovery), 0.13 (edge sharpness), 0.06 (naturalness), with Fleiss κ negative on all three scales.
The divergence is interpretable rather than random. QA testers scored naturalness as the absence of artifacts, and artifact generation is SeedVR2’s characteristic failure. Distribution reviewers scored it as fidelity to the source, and infidelity is Topaz Wonder 3’s characteristic failure. Both readings are defensible under the scale definition used here, which indicates the definition is underspecified for panels that mix roles.
3.5 Inter-rater agreement
The panel, the interface, the categories, and the pair count were unchanged between sessions. The variable that changed is the size of the difference under evaluation. When two outputs differ by less than a rater can resolve on a single pair, ratings on undecidable pairs approach random selection among the three options and agreement statistics fall toward zero by construction.
The agreement collapse is therefore treated as a measurement of proximity, not as a data-quality failure. It is also the reason this session is labelled directional: under the internal protocol, a verified label requires agreement above threshold, and no aggregate reported here meets it.
4. Qualitative Evaluator Reports
Nineteen structured summaries were submitted by 10 of the 13 evaluators. Attribution is by role; individual evaluators are not named.
SeedVR2
Reported strengths. Edge definition and recovery of structure present in the source; closer preservation of facial proportions and likeness; strongest performance on materials, metal, and object textures; reproduction of fine repeating textures such as grass without introducing new patterns; no license cost.
Reported failure modes. Over-sharpening to the point that the result reads as rendered rather than photographed; reconstruction of detail absent from the source, including hair added to a face and an incorrectly reconstructed skin mark; higher error rate on small text and numerals; occasional synthetic skin texture.
Topaz Wonder 3
Reported strengths. More coherent skin texture at portrait scale, better retention of small text and numerals in the documented cases, smooth colour transitions, appropriate sharpening on wide and product shots, and fine-grained parameter control.
Reported failure modes. Visible alteration of facial features; skin rendered as a plastic-looking surface; elevated sharpening and contrast producing a filtered appearance; repeating patterns on grass and uniform textures.
Two summaries state the aggregate position directly.
Practically identical to Topaz in quality, but given that it is free, I would pick it for real work.
— QA tester, on SeedVR2
The best upscaler right now, but the big drawback is that it is paid, and it is not much better than SeedVR2.
— Art director, on Topaz Wonder 3
4.1 Small text and numerals
The documented instances are specific: on a cable label, SeedVR2 reconstructed a “2” as “3”, and a footnote mark as “0”; Topaz Wonder 3 preserved both. Evaluators also reported that both models degrade text that was already illegible or incorrectly rendered in the source. The sample of documented text cases is small and the effect is not separated statistically, so this is a flag for verification rather than a measured advantage.
5. Limitations
- Agreement below threshold. All aggregate figures are directional. Reported confidence intervals describe sampling variation, not rater reliability, which is separately low.
- Single session, single configuration. Topaz Wonder 3 exposes tunable parameters; the session used one configuration per model, and parameter tuning could shift the result on any scale.
- Subjective scales only. No reference-based metrics (PSNR, SSIM, LPIPS) were computed. The evaluation measures preference by trained reviewers on production content, not distortion against ground truth.
- Three ratings per pair. With three raters and three response options, per-pair majority is decided by two votes on most pairs.
- Scale definition. The between-role divergence on naturalness (Section 3.4) indicates that the naturalness scale admits at least two consistent readings, artifact absence and source fidelity, which are not separated in the instrument.
- Content scope. Source images come from Everypixel production stock workflows. Results should not be extrapolated to medical, scientific, forensic, or archival restoration.
6. Practical Implications
Model selection
Default to SeedVR2 for general upscaling. It matches Topaz Wonder 3 on the majority of content, leads on edge sharpness and on materials and textures, preserves likeness more closely, and carries no license cost. The measured quality difference does not support a paid step in a high-volume pipeline.
Route portrait work to Topaz Wonder 3 where skin must read as unprocessed. This is the only category in which the panel produced a clear preference, and it runs against the overall result. Where fidelity to the subject’s actual features matters more than skin appearance, the preference inverts back to SeedVR2.
Verify any output carrying legible text, regardless of model. Topaz Wonder 3 degraded less in the documented cases, but neither model is safe on labels, packaging, interface captures, or documents without a check.
A 0.55 overall probability is not grounds for migrating an existing workflow. Migration cost is certain, the quality delta is not visible on most images, and parameter control in Topaz Wonder 3 has value for teams that tune per batch.
7. FAQ
Which AI upscaler is better in 2026, SeedVR2 or Topaz Wonder 3?
Neither, on most content. In this July 2026 evaluation of 264 image pairs, 54–60% of comparisons were ties and the overall Bradley-Terry probability that SeedVR2 is preferred was 0.55 (95% CI 0.53–0.58). SeedVR2 leads on edge sharpness (P = 0.60); Topaz Wonder 3 leads on faces. Since SeedVR2 carries no license cost, it is the better-supported default.
Is Topaz Wonder 3 worth paying for if SeedVR2 is free?
For general-purpose upscaling, this data does not support a quality-based case. For portrait work it does: Topaz Wonder 3 was preferred on naturalness for 52% of face pairs against 14%. Its parameter control is also a functional difference not captured by a pairwise preference test.
Which upscaler is better for portraits and faces?
Topaz Wonder 3 on naturalness (52% against 14% of face pairs, July 2026). The trade-off is fidelity: evaluators reported that Topaz Wonder 3 alters facial features more, while SeedVR2 preserves proportions and likeness but can reconstruct details that were not present in the source.
Can either model handle small text and numbers?
Not reliably. Text and signage produced the highest tie rates in the session (64–68%), and evaluators documented SeedVR2 changing a “20” to a “30” and a footnote mark to a “0” where Topaz Wonder 3 preserved both. Verify any upscaled image carrying legible text.
Why is evaluator agreement so low in this benchmark?
Because the models are close. Gwet’s AC1 was 0.06–0.12 here against 0.42–0.71 for the same panel in June 2026. When the difference between two outputs falls below what a trained rater can resolve on a single pair, ratings on undecidable pairs approach random selection and agreement falls toward zero. The result is published as directional for that reason.
Why report Gwet’s AC1 instead of Fleiss kappa?
Because ties account for roughly 60% of outcomes. Under a skewed outcome distribution the Fleiss κ expected-agreement term inflates and κ collapses toward zero regardless of actual concordance (Feinstein & Cicchetti, 1990). AC1 is stable under skew (Gwet, 2002). Fleiss κ is reported in this study as well and reads 0.00–0.05, which is the artifact rather than a measurement.
How large was this evaluation?
264 image pairs (132 real photographs, 132 AI-generated) across six categories, three independent ratings per pair on three quality scales, 791 completed comparisons. Thirteen evaluators: 4 QA testers, 4 art directors, 5 distribution reviewers. Session ff12fe31, closed July 2026.
8. Data Availability
Per-comparison records (264 pairs × rater × scale, with category, source type, and evaluator role); aggregate statistics, including all confidence intervals; and the 19 structured evaluator summaries are retained under session ff12fe31 and are available on request for research purposes.
About Everypixel
Everypixel runs systematic benchmarks of the AI image and video models used in production content workflows. Evaluations use domain-specific reviewer panels rather than crowdsourced annotation, so results reflect commercial visual content standards. Research is published at research.everypixel.com.
Related: SeedVR2 vs FLUX.2 [klein] 9B, the June 2026 upscaler benchmark run by the same panel on the same 264-pair structure.
Author: Everypixel Production Team · Last reviewed: August 2026 · Confidence level: directional
References
- ByteDance Seed. (2026). SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training. arXiv:2506.05301. https://arxiv.org/abs/2506.05301
- Bradley, R.A., & Terry, M.E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika, 39(3/4), 324–345. https://doi.org/10.2307/2334029
- Feinstein, A.R., & Cicchetti, D.V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543–549. https://doi.org/10.1016/0895-4356(90)90158-L
- Gwet, K.L. (2002). Kappa statistic is not satisfactory for assessing the extent of agreement between raters. Statistical Methods for Inter-Rater Reliability Assessment, 1, 1–6. https://agreestat.com/papers/kappa_statistic_is_not_satisfactory.pdf
Cite this article
<blockquote cite="https://research.everypixel.com/seedvr2-vs-topaz-wonder-3-a-264-pair-blind-evaluation-of-two-image-upscalers/"> <p>In the research, 13 evaluators compared SeedVR2 and Topaz Wonder 3 on 264 image pairs and tied 54-60% of them. SeedVR2 leads on edge sharpness, Topaz Wonder 3 wins faces, and the free model is the defensible default.</p> <footer>— <a href="https://research.everypixel.com/seedvr2-vs-topaz-wonder-3-a-264-pair-blind-evaluation-of-two-image-upscalers/">SeedVR2 vs Topaz Wonder 3: A 264-Pair Blind Evaluation of Two Image Upscalers</a>, Everypixel Research, August 2026</footer> </blockquote>
Everypixel Research. (2026). SeedVR2 vs Topaz Wonder 3: A 264-Pair Blind Evaluation of Two Image Upscalers. research.everypixel.com. https://research.everypixel.com/seedvr2-vs-topaz-wonder-3-a-264-pair-blind-evaluation-of-two-image-upscalers/