GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท gpt-4-1

GPT-4.1

vgpt-4.1-2025-04-14

OpenAI

Modelgeneral-purposeflagshipproduction-readymultimodal
85
Strong
About This Model

LEGACY: retired from ChatGPT 2026-02-13 but still available in the API with no announced shutdown (as of 2026-07-09). Previous-generation general-purpose GPT-4.1 model with a 1,047,576-token context window. OpenAI recommends GPT-5.x models for new work.

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+

Strong general-purpose performance with good balance across coding, reasoning, and knowledge tasks. Flagship model for most production use cases.

task accuracy code

Industry-standard coding benchmarks

Evidence
HumanEval Benchmark โ€” 48.1% pass rate
MBPP Benchmark โ€” 62% on mostly basic programming problems
highVerified: 2026-07-09
task accuracy reasoning

Mathematical and scientific reasoning benchmarks

Evidence
MATH Benchmark โ€” 68% on mathematical reasoning tasks
GPQA โ€” 52% on graduate-level reasoning
highVerified: 2026-07-09
task accuracy general

Crowdsourced comparisons and comprehensive knowledge testing

Evidence
MMLU Benchmark โ€” 66.3% on multitask language understanding
LMSYS Chatbot Arena โ€” 1250 ELO (Strong mid-tier performance)
highVerified: 2026-07-09
output consistency

Internal testing with repeated prompts

Evidence
OpenAI Internal Testing โ€” Strong consistency across temperature settings
highVerified: 2026-07-09
latency p50

Median latency for API requests

Evidence
OpenAI Documentation โ€” Typical response time ~1.2s
highVerified: 2026-07-09
latency p95

95th percentile response time

Evidence
Community benchmarking โ€” p95 latency ~2.4s
highVerified: 2026-07-09
context window

Official specification from provider

Evidence
OpenAI Model Page: gpt-4.1 โ€” 1,047,576 token context window; 32,768 max output tokens
highVerified: 2026-07-09
uptime

Historical uptime data from official status page

Evidence
OpenAI Status Page โ€” 99.9% uptime (last 90 days)
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+

Strong security posture with comprehensive safety systems. Robust protection against adversarial attacks.

prompt injection resistance

Testing against OWASP LLM01 prompt injection attacks

Evidence
OpenAI Safety Testing โ€” Strong resistance to prompt injection attacks
highVerified: 2026-07-09
jailbreak resistance

Testing against adversarial prompt datasets

Evidence
OpenAI Safety Evaluations โ€” Robust safety mechanisms
highVerified: 2026-07-09
data leakage prevention

Analysis of privacy policies

Evidence
OpenAI Privacy Policy โ€” API data not used for training by default
mediumVerified: 2026-07-09
output safety

Safety testing across harmful content categories

Evidence
OpenAI Safety Benchmarks โ€” Comprehensive safety systems
highVerified: 2026-07-09
api security

Review of API security features

Evidence
OpenAI API Documentation โ€” API key authentication, HTTPS, rate limiting
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+

Standard enterprise privacy practices with SOC 2 Type II certification. 30-day retention period.

data residency

Review of enterprise documentation

Evidence
OpenAI Documentation โ€” US-based infrastructure
highVerified: 2026-07-09
training data optout

Analysis of privacy policy

Evidence
OpenAI Privacy Policy โ€” API data not used for training by default
highVerified: 2026-07-09
data retention

Review of terms of service

Evidence
OpenAI Terms of Service โ€” API data retained for 30 days
highVerified: 2026-07-09
pii handling

Review of data protection capabilities

Evidence
OpenAI Privacy Documentation โ€” Customer responsible for PII redaction
mediumVerified: 2026-07-09
compliance certifications

Verification of compliance certifications

Evidence
OpenAI Trust Portal โ€” SOC 2 Type II, GDPR compliant
highVerified: 2026-07-09
zero data retention

Review of data handling practices

Evidence
OpenAI API Documentation โ€” 30-day retention for abuse monitoring
highVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+

Good transparency with solid explainability. Lower hallucination rate than smaller models. Comprehensive safety systems.

explainability

Evaluation of reasoning transparency

Evidence
Model Behavior โ€” Good explanations and reasoning
mediumVerified: 2026-07-09
hallucination rate

Testing on factual QA datasets

Evidence
SimpleQA Benchmark โ€” Good factual accuracy
mediumVerified: 2026-07-09
bias fairness

Evaluation on bias benchmarks

Evidence
OpenAI Safety Report โ€” Regular bias testing and mitigation
mediumVerified: 2026-07-09
uncertainty quantification

Qualitative assessment of confidence expression

Evidence
Model Behavior โ€” Good uncertainty expression
mediumVerified: 2026-07-09
model card quality

Review of documentation completeness

Evidence
OpenAI Model Documentation โ€” Comprehensive documentation with benchmarks
highVerified: 2026-07-09
training data transparency

Review of public disclosures

Evidence
OpenAI Public Statements โ€” General description provided
mediumVerified: 2026-07-09
guardrails

Analysis of safety mechanisms

Evidence
OpenAI Safety Systems โ€” Comprehensive safety guardrails
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+

Excellent operational maturity with industry-leading ecosystem and developer experience.

api design quality

Review of API design

Evidence
OpenAI API Documentation โ€” Well-designed RESTful API with comprehensive features
highVerified: 2026-07-09
sdk quality

Review of SDK quality

Evidence
OpenAI SDKs โ€” High-quality SDKs for Python, Node.js
highVerified: 2026-07-09
versioning policy

Review of versioning approach

Evidence
OpenAI API Versioning โ€” Clear versioning with deprecation notices
OpenAI: Retiring GPT-4o and older models โ€” GPT-4.1 retired from ChatGPT 2026-02-13; API access continues with no announced shutdown (not on the API deprecations list as of 2026-07-09)
highVerified: 2026-07-09
monitoring observability

Review of monitoring tools

Evidence
OpenAI Dashboard โ€” Comprehensive usage dashboard
mediumVerified: 2026-07-09
support quality

Assessment of support channels

Evidence
OpenAI Support โ€” Excellent support and documentation
highVerified: 2026-07-09
ecosystem maturity

Analysis of integrations

Evidence
GitHub Ecosystem โ€” Extremely mature ecosystem
highVerified: 2026-07-09
license terms

Review of licensing

Evidence
OpenAI Terms of Service โ€” Clear commercial terms
highVerified: 2026-07-09
Strengths
  • +Strong general-purpose performance (66.3% MMLU)
  • +Good balance of quality and speed (~1.2s p50)
  • +Very large context window (1,047,576 tokens) for document processing
  • +Mature ecosystem with extensive integrations
  • +Reliable uptime and infrastructure (99.9%)
  • +Comprehensive safety and security features
Limitations
  • !Moderate coding performance (48.1% HumanEval)
  • !30-day data retention period
  • !Not HIPAA eligible
  • !Limited regional data residency options
  • !Higher pricing than smaller models
  • !Training data transparency limited
  • !LEGACY: retired from ChatGPT 2026-02-13; API continues but OpenAI recommends GPT-5.x for new work
Metadata
pricing
input: $2.00 per 1M tokens
output: $8.00 per 1M tokens
notes: Cached input $0.50 per 1M. Confirmed on official model page 2026-07-09; no longer listed on OpenAI's main pricing page.
last verified: 2026-07-09
context window: 1047576
max output: 32768
languages
0: English
1: Spanish
2: French
3: German
4: Italian
5: Portuguese
6: Japanese
7: Korean
8: Chinese
9: Arabic
10: Hindi
11: Russian
12: Dutch
modalities
0: text
1: image (input)
api endpoint: https://api.openai.com/v1/chat/completions
open source: false
architecture: Transformer-based with multimodal capabilities
parameters: Not disclosed (large)

Use Case Ratings

code generation

Good coding capabilities for typical development tasks. 48.1% HumanEval suitable for standard programming.

customer support

Excellent for customer support with strong conversational abilities and good response times.

content creation

Strong content creation with natural language and good creativity.

data analysis

Good for data analysis and business intelligence tasks.

research assistant

Strong research capabilities with good knowledge base (66.3% MMLU).

legal compliance

Adequate for legal document analysis but requires human oversight.

healthcare

Not HIPAA eligible. Limited use for healthcare applications.

financial analysis

Good for financial analysis and reporting tasks.

education

Excellent for educational applications and tutoring.

creative writing

Strong creative writing with natural storytelling abilities.