GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท claude-opus-4-7

Claude Opus 4.7

v20260416

Anthropic

Modelcodingreasoningvisionenterprise
91
Exceptional
About This Model

Previous-generation Opus flagship, superseded by Opus 4.8. 64.3% SWE-Bench Pro and 94.2% GPQA Diamond at launch. First Claude with high-resolution vision (2576px long edge, pixel-accurate coordinates), task budgets (beta), and the xhigh effort level.

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+

Was Anthropic's most capable model at launch (2026-04-16); now the previous-generation Opus behind Opus 4.8. Remains a strong, fully supported flagship-class choice, especially for vision-heavy workloads.

task accuracy code

Industry-standard coding benchmarks measuring real-world software engineering tasks

Evidence
SWE-Bench Pro โ€” 64.3% on SWE-Bench Pro at launch
Anthropic Launch Announcement โ€” Strong long-horizon agentic coding; meaningfully better bug-finding recall and precision than prior Opus models
highVerified: 2026-07-09
task accuracy reasoning

Graduate and PhD-level reasoning benchmarks evaluated with adaptive thinking at high effort

Evidence
GPQA Diamond โ€” 94.2% on PhD-level science questions
highVerified: 2026-07-09
task accuracy general

Comprehensive knowledge and multimodal testing, including high-resolution screenshot and document understanding

Evidence
Anthropic Launch Announcement โ€” First Claude with high-resolution vision (2576px long edge); gains in knowledge work, memory, and vision-heavy tasks
highVerified: 2026-07-09
output consistency

Repeated-run consistency testing across effort levels; structured extraction pipelines

Evidence
Anthropic Launch Announcement โ€” More literal instruction following and stricter effort adherence yield more predictable pipeline behavior
highVerified: 2026-07-09
latency p50

Median latency for API requests with standard prompt sizes

Evidence
Community benchmarking โ€” Typical response time ~2.8s for standard prompts at default effort
mediumVerified: 2026-07-09
latency p95

95th percentile response time across diverse workloads

Evidence
Community benchmarking โ€” p95 ~6.5s; higher at xhigh/max effort
mediumVerified: 2026-07-09
context window

Official specification from provider

Evidence
Anthropic API Documentation โ€” 1M token context window, 128K max output tokens
highVerified: 2026-07-09
uptime

Historical uptime data from official status page

Evidence
Anthropic Status Page โ€” Claude API uptime 99.57% (last 90 days)
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+

Introduced real-time cybersecurity safeguards to the Opus line. Strong overall posture carried forward into Opus 4.8.

prompt injection resistance

Testing against OWASP LLM01 prompt injection attacks

Evidence
Anthropic Launch Announcement โ€” Improved resistance to injected instructions; more literal instruction-following reduces susceptibility to embedded directives
highVerified: 2026-07-09
jailbreak resistance

Testing against adversarial prompt datasets

Evidence
Anthropic Constitutional AI โ€” Constitutional AI alignment with refined refusal calibration
highVerified: 2026-07-09
data leakage prevention

Analysis of privacy policies and data handling practices

Evidence
Anthropic Privacy Statement โ€” Training opt-out by default for API; no training on user data without consent
mediumVerified: 2026-07-09
output safety

Comprehensive safety testing across harmful content categories

Evidence
Anthropic Launch Announcement โ€” Introduced real-time cybersecurity safeguards for prohibited and high-risk topics
highVerified: 2026-07-09
api security

Review of API security features and best practices

Evidence
Anthropic API Documentation โ€” API key and OAuth authentication, HTTPS only, rate limiting, workspace scoping
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+

Standard Anthropic enterprise compliance posture: SOC 2 Type II, GDPR, HIPAA-eligible, training opt-out by default for API traffic.

data residency

Review of enterprise documentation and privacy policies

Evidence
Anthropic Enterprise Documentation โ€” Data residency options for US and EU enterprise customers
highVerified: 2026-07-09
training data optout

Analysis of privacy policy and data usage terms

Evidence
Anthropic Privacy Policy โ€” Training opt-out by default for API usage
highVerified: 2026-07-09
data retention

Review of terms of service and data retention policies

Evidence
Anthropic Trust Center โ€” Zero data retention agreements available for eligible API customers
highVerified: 2026-07-09
pii handling

Review of data protection capabilities and customer responsibilities

Evidence
Anthropic Privacy Documentation โ€” Customer responsible for PII redaction; provider-side safeguards for incidental PII
mediumVerified: 2026-07-09
compliance certifications

Verification of compliance certifications and audit reports

Evidence
Anthropic Trust Center โ€” SOC 2 Type II, GDPR compliant, HIPAA eligible
highVerified: 2026-07-09
zero data retention

Review of data handling practices and trust center documentation

Evidence
Anthropic Trust Center โ€” Zero data retention configuration available; no training on API data by default
highVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+

Strong documentation and guardrails. Thinking content omitted by default reduces out-of-the-box reasoning visibility; opt in to summarized display if reasoning is surfaced to users.

explainability

Evaluation of reasoning transparency and explanation capabilities

Evidence
Anthropic Adaptive Thinking Documentation โ€” Adaptive thinking with effort control; thinking content omitted by default โ€” summarized display is opt-in
mediumVerified: 2026-07-09
hallucination rate

Testing on factual QA datasets and document-fidelity evaluations

Evidence
Anthropic Launch Announcement โ€” Improved factual accuracy; visually verifies its own output in knowledge-work tasks
mediumVerified: 2026-07-09
bias fairness

Evaluation on bias benchmarks and diverse demographic testing

Evidence
Anthropic Responsible Scaling Policy โ€” Regular bias testing and mitigation under the Responsible Scaling Policy
mediumVerified: 2026-07-09
uncertainty quantification

Qualitative assessment of confidence expression in outputs

Evidence
Anthropic Launch Announcement โ€” More literal, precise behavior; expresses uncertainty rather than inferring unrequested intent
mediumVerified: 2026-07-09
model card quality

Review of documentation completeness and clarity

Evidence
Anthropic Model Documentation โ€” Comprehensive model card with capabilities, limitations, benchmarks, and migration guidance
highVerified: 2026-07-09
training data transparency

Review of public disclosures about training data

Evidence
Anthropic Public Statements โ€” General description provided, detailed sources not disclosed
mediumVerified: 2026-07-09
guardrails

Analysis of built-in safety mechanisms

Evidence
Constitutional AI โ€” Constitutional AI safety guardrails with real-time cybersecurity safeguards
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+

Mature operational profile. Superseded by Opus 4.8 as the flagship Opus, but remains fully supported at the same $5/$25 price; upgrade to 4.8 is a drop-in model-ID swap.

api design quality

Review of API design, consistency, and feature completeness

Evidence
Anthropic Migration Guide โ€” Introduced the xhigh effort level and task budgets (beta); adaptive thinking only โ€” budget_tokens and temperature/top_p/top_k removed
highVerified: 2026-07-09
sdk quality

Review of SDK quality, documentation, and maintenance

Evidence
Anthropic SDKs โ€” Official SDKs (Python, TypeScript, Java, Go, Ruby, C#, PHP) with day-one support
highVerified: 2026-07-09
versioning policy

Review of versioning policy and historical practices

Evidence
Anthropic API Versioning โ€” Clear versioning with advance deprecation notice; claude-opus-4-7 alias remains active after Opus 4.8 launch
Anthropic Model Deprecations โ€” Active; tentative retirement not sooner than April 16, 2027
highVerified: 2026-07-09
monitoring observability

Review of available monitoring tools and metrics

Evidence
Anthropic Console โ€” Usage dashboard with metrics, cost tracking, and workspace controls
mediumVerified: 2026-07-09
support quality

Assessment of documentation, community, and support responsiveness

Evidence
Anthropic Support โ€” Email support, developer community, comprehensive docs and migration guides
highVerified: 2026-07-09
ecosystem maturity

Analysis of third-party integrations and availability surfaces

Evidence
Anthropic Launch Announcement โ€” Available on the Anthropic API and major cloud platforms; broad tooling support
highVerified: 2026-07-09
license terms

Review of licensing terms and restrictions

Evidence
Anthropic Commercial Terms โ€” Standard commercial terms; enterprise agreements available
highVerified: 2026-07-09
Strengths
  • +64.3% SWE-Bench Pro and 94.2% GPQA Diamond at launch
  • +First Claude with high-resolution vision: 2576px long edge with pixel-accurate coordinates
  • +Introduced the xhigh effort level and task budgets (beta) for agentic token control
  • +1M token context window with 128K max output at $5/$25
  • +More literal, predictable instruction following for tuned pipelines
  • +Strong compliance posture: SOC 2 Type II, GDPR, HIPAA-eligible, training opt-out by default
Limitations
  • !Superseded by Opus 4.8 as the current flagship Opus (same price, drop-in upgrade)
  • !Adaptive thinking only โ€” budget_tokens and temperature/top_p/top_k return 400
  • !Thinking content omitted by default; summarized display requires opt-in
  • !Full-resolution images can use up to ~3x more image tokens than prior models
  • !Reaches for tools and subagents less often than Opus 4.6 without explicit prompting
Metadata
pricing
input: $5.00 per 1M tokens
output: $25.00 per 1M tokens
notes: 1M context at standard API pricing with no long-context premium. Batch API 50% discount, prompt caching savings apply. No fast-mode variant. Confirmed unchanged at $5/$25 as of 2026-07-09.
last verified: 2026-07-09
context window: 1000000
max output: 128000
languages
0: English
1: Spanish
2: French
3: German
4: Italian
5: Portuguese
6: Japanese
7: Korean
8: Chinese
9: Arabic
10: Hindi
modalities
0: text
1: image (input, high-resolution)
2: document
3: computer-use
api endpoint: https://api.anthropic.com/v1/messages
api model id: claude-opus-4-7
open source: false
architecture: Transformer-based with Constitutional AI alignment; adaptive thinking only with effort parameter introducing xhigh; high-resolution vision
parameters: Not disclosed
knowledge cutoff: January 2026 (reliable knowledge and training data cutoff)
release date: 2026-04-16

Use Case Ratings

code generation

64.3% SWE-Bench Pro with strong long-horizon agentic coding and improved bug-finding. Superseded by Opus 4.8 at the same price โ€” prefer 4.8 for new builds.

customer support

High quality but more clipped, direct tone than 4.8; Sonnet/Haiku tiers are more cost-effective for routine volume.

content creation

Strong long-form output, though more terse and less warm than Opus 4.8 by default; style is prompt-tunable.

data analysis

Excellent analytical depth; high-resolution vision enables pixel-level chart and figure transcription.

research assistant

Strong deep research with 1M context and improved file-based memory; Opus 4.8 improves further on long-horizon coherence.

legal compliance

Strong privacy posture (SOC 2 Type II, GDPR, HIPAA-eligible) and literal instruction following suited to compliance pipelines.

healthcare

HIPAA eligible with training opt-out by default; high-resolution vision aids medical document and chart understanding.

financial analysis

Excellent quantitative reasoning; pixel-accurate chart reading and 1M context handle full filings and figures.

education

Clear, precise explanations with effort-adjustable depth; more literal style benefits structured curricula.

creative writing

Capable but more clipped and direct than Opus 4.8's warmer voice; no sampling parameters, so variety must be prompted.