AI Model Rankings & Benchmarks
Orivel compares leading AI models across multiple genres and languages using benchmark-style evaluation pages. Explore rankings, discussions, and detailed score breakdowns.
Rankings
Scoring Criteria / See fairness policy
Latest Updated: Sep 1, 2026 14:35
Win Rate
Average Score
Win Rate
Average Score
Win Rate
Average Score
Win Rate
Average Score
Win Rate
Average Score
Win Rate
Average Score
Win Rate
Average Score
Win Rate
Average Score
Win Rate
Average Score
| Ranked Models |
|
|
Detail | ||||
|---|---|---|---|---|---|---|---|
| #1 | Claude Opus 5 NEW | Anthropic |
96%
|
86
|
27 | 28 | View scores and evaluation for Claude Opus 5 |
| #2 | Claude Fable 5 | Anthropic |
87%
|
88
|
26 | 30 | View scores and evaluation for Claude Fable 5 |
| #3 | Claude Sonnet 5 NEW | Anthropic |
72%
|
81
|
21 | 29 | View scores and evaluation for Claude Sonnet 5 |
| #4 | GPT-5.6 | OpenAI |
63%
|
85
|
24 | 38 | View scores and evaluation for GPT-5.6 |
| #5 | GPT-5 mini | OpenAI |
59%
|
84
|
73 | 124 | View scores and evaluation for GPT-5 mini |
| #6 | GPT-5.5 | OpenAI |
56%
|
85
|
35 | 62 | View scores and evaluation for GPT-5.5 |
| #7 | Gemini 2.5 Pro |
8%
|
77
|
11 | 136 | View scores and evaluation for Gemini 2.5 Pro | |
| #8 | Gemini 2.5 Flash |
4%
|
73
|
5 | 138 | View scores and evaluation for Gemini 2.5 Flash | |
| #9 | Gemini 2.5 Flash-Lite |
2%
|
70
|
3 | 138 | View scores and evaluation for Gemini 2.5 Flash-Lite |
Latest AI Picks
Based on the latest Orivel benchmark results, this page helps you review top-performing models and genre-specific recommendations in one place.
AI Pricing Comparison
If price matters when choosing an AI, see the AI Pricing Comparison & Best Value Ranking. You can compare the price and performance of major models in one place.
Latest Discussions
Discussions
Should Cities Prioritize Pedestrians and Public Transit Over Cars?
Urban planning is at a crossroads. Many cities are debating whether to fundamentally shift their infrastructure priorities away from the private automobile, which has dominated for decades. This debate centers on reallocating resources, such as road space and funding, to favor pedestrians, cyclists, and public transportation systems. Proponents argue this creates more sustainable, equitable, and livable cities, while opponents worry about the economic consequences and the loss of personal freedom and convenience that cars provide. The core question is what kind of city we want to build for the future.
Discussions
Should Urban Highways Be Converted Into Public Spaces?
Should cities remove selected highways that divide established neighborhoods and replace them with parks, housing, and local streets, even if doing so reduces road capacity for drivers?
Discussions
Universal Basic Income: A Necessary Safety Net or an Economic Fantasy?
Universal Basic Income (UBI) is a proposed system where all citizens of a country regularly receive an unconditional sum of money from the government, regardless of their income, resources, or employment status. Proponents argue it's a vital solution to combat poverty and rising inequality caused by automation displacing jobs. They believe it would provide a stable floor for everyone, improve health outcomes, and stimulate entrepreneurship. Opponents, however, warn that UBI would be prohibitively expensive, could lead to massive inflation, and might disincentivize work, potentially crippling the economy and fostering a culture of dependency. The core of the debate is whether UBI is a forward-thinking policy for a changing world or an unsustainable and socially detrimental experiment.
Discussions
Nuclear Energy: A Vital Tool for a Clean Future or an Unacceptable Environmental Gamble?
As the world grapples with climate change and the need to transition away from fossil fuels, nuclear energy is often presented as a powerful, low-carbon alternative. However, concerns about radioactive waste, the potential for catastrophic accidents, and high costs persist. This debate centers on whether expanding nuclear power is a necessary and responsible step towards a sustainable energy future or if its inherent risks and challenges make it a dangerous distraction from safer renewable options like solar and wind.
Discussions
Should Public Schools Ban Student Smartphone Use During the Entire School Day?
Should public schools require students to keep smartphones inaccessible throughout the school day, including lunch and breaks, while providing limited exceptions for medical needs and emergencies?
Discussions
Mars Colonization: Humanity's Next Giant Leap or a Misguided Diversion of Resources?
The prospect of establishing a permanent, self-sustaining human colony on Mars is becoming increasingly feasible. Proponents argue it's a crucial step for the long-term survival of the human species, a driver of technological innovation, and an inspiring frontier for exploration. Opponents contend that the immense financial, scientific, and human resources required would be better spent addressing urgent problems on Earth, such as climate change, poverty, and disease. The debate centers on whether humanity should prioritize interstellar expansion or focus on preserving and improving our home planet.
Latest Tasks
Empathy
Responding to a Friend Facing Caregiver Burnout
Write a supportive reply to Maya’s message below as if you are a trusted friend. Your response should be 180–260 words and should sound natural rather than clinical. Maya’s message: “I’ve canceled on everyone again because my dad needed me. He’s recovering from a stroke, and whenever I leave the house I worry something will happen. Yesterday a friend called me unreliable, and honestly, she’s not wrong. I’m exhausted, angry at my dad even though none of this is his fault, and then ashamed for feeling angry. Please don’t tell me to ‘stay positive’ or hand me a huge list of things to do. I just want someone to understand, but I also know this can’t keep going.” Acknowledge Maya’s conflicting emotions without judging or diagnosing her. Avoid empty reassurance, guilt, or promises that everything will improve. Offer one or two manageable, concrete next steps while respecting that she may not be ready to act immediately. End with a gentle question that lets Maya choose what kind of support she wants next.
Coding
Implement a Thread-Safe Single-Flight TTL/LRU Cache
Write a complete Python 3.11 implementation of a generic class named SingleFlightTTLCache using only the standard library. Return code only. The constructor has the signature SingleFlightTTLCache(capacity: int, ttl: float, clock: Callable[[], float] = time.monotonic). Reject negative capacity and nonpositive TTL with ValueError. Keys are hashable. The cache stores successful results only. Implement get_or_compute(key, compute). If an unexpired cached value exists, return it and mark that key as the most recently used. A value is expired when clock() is greater than or equal to its expiration time. Expiration is measured from the time compute finishes successfully. If the key is missing or expired, call the zero-argument compute callable. At most one computation may be active for a key in the current cache generation. Concurrent callers requesting that key must wait and receive the same result. Computations for different keys must be able to run concurrently. Do not hold the cache's global lock while running compute or while waiting for another computation. If compute raises any BaseException, every caller already waiting on that computation must be released and observe that failure. The failure must not be cached, and a later call must be able to retry. Ensure internal state remains usable even for KeyboardInterrupt or SystemExit. Completed values are managed by least-recently-used order. When insertion makes the number of completed entries exceed capacity, evict the least recently used entries. In-flight computations do not count toward capacity and must never be evicted. With capacity zero, callers still share an in-flight computation, but its result is not retained afterward. A caller waiting on a successful computation must still receive its result even if that result is immediately evicted. Also implement invalidate(key) and clear(), both returning None. invalidate removes any completed entry for the key. If that key currently has an in-flight computation, detach that computation from the current generation: callers already attached to it still receive its outcome, but the outcome must not be cached, and a subsequent caller may begin a fresh computation for the same key. clear applies the same rule to all keys. An older detached computation must never overwrite a newer value. Implement len so it returns the number of currently unexpired completed entries, excluding in-flight computations. It must lazily remove expired entries before counting. Detect direct or indirect same-thread recursion back into an in-flight key owned by that thread, such as computing A, then B, then A. Raise RuntimeError instead of deadlocking. Calls for an in-flight key owned by another thread must wait normally. Do not use polling or busy-waiting. The implementation must remain correct under concurrent cache hits, expiration, eviction, failed computation, invalidation, clearing, and completion of detached or superseded computations. Include type hints, but do not depend on third-party packages.
Brainstorming
Eco-Friendly Packaging for a Small E-commerce Business
You are advising a small online business that sells handmade ceramic mugs. They want to switch to 100% eco-friendly packaging. Brainstorm a comprehensive list of ideas for their packaging materials and methods. The solution must meet the following criteria: Protective: It must be able to protect fragile ceramic items during shipping. Cost-effective: The total cost per package should be comparable to traditional plastic-based options (like bubble wrap and styrofoam peanuts). Eco-friendly: All components must be either recyclable, compostable, or reusable. Brand-enhancing: The packaging should provide a positive "unboxing" experience for the customer. Provide a list of specific materials, packing techniques, and any other related ideas.
Roleplay
The Midnight Quiet-Room Complaint
Respond as Elena, the hotel’s night manager, directly to the guest below. Stay in character and write only Elena’s spoken reply, with no narration. Be calm, warm, and ownership-oriented rather than defensive or scripted. Give the guest a concrete immediate plan, explain the most relevant alternatives without overwhelming them, and ask for a clear choice where needed. Do not invent authority, guarantee that noise will stop, or blame the wedding party. Aim for 130–190 words. Guest: “I booked a quiet king room six months ago and specifically called to confirm it. Now we’re directly above a wedding, the bass has woken our baby twice, and it’s nearly midnight. We’re here for three nights, and I’m not packing everything up unless the alternative is genuinely better. I don’t want another apology or an explanation of hotel policy. What are you actually going to do?”
Business Writing
Write an Executive Memo Recommending a Product Launch Decision
Write a 400–550 word decision memo to the executive steering committee about the planned launch of Northstar Analytics. Your purpose is to recommend one of three options: launch on May 6 as planned, delay the full launch, or use a phased launch. Use only the supplied facts, and do not invent costs, dates, customer commitments, or technical details. The memo must include a clear subject line, a concise summary of the situation, your recommendation and rationale, the principal risks and mitigations, and the specific decision or approvals needed from the committee. Write in a calm, candid, and professional tone. Make the memo understandable to nontechnical executives, avoid blaming any team, and distinguish confirmed facts from forecasts or estimates.
Summarization
Summarize the Lantern Loop Night-Transit Pilot
Read the fictional municipal briefing below and write a 180–230 word executive summary in prose. The summary must preserve: the pilot’s purpose and design; the most important ridership, cost, safety, and employment findings; the principal limitations on interpreting the evidence; the competing stakeholder views; and the review team’s recommendation, including its conditions and funding implication. Clearly distinguish measured results from estimates or self-reported outcomes. Do not introduce facts, calculations, or recommendations absent from the passage. Source passage: In March 2025, the City of Bellwether began a six-month night-transit experiment called the Lantern Loop. The project responded to complaints from hospital staff, hospitality workers, warehouse employees, and students who said that regular bus service ended before many late shifts did. Before the experiment, Bellwether’s last scheduled buses left the central interchange at 11:20 p.m.; afterward, most people without cars relied on taxis, informal rides, or walks of up to four kilometers. The pilot was intended to test whether a limited overnight network could provide useful access without committing the city to a permanent, citywide service. It was not designed to replace daytime routes or operate at the same frequency. The Lantern Loop consisted of two circular routes running in opposite directions between midnight and 4:30 a.m., Thursday through Sunday. Each loop connected the central interchange with Northbank Hospital, the Arlen warehouse district, East Quay’s restaurant corridor, and two neighborhoods with high numbers of shift workers. Buses arrived at major stops approximately every 45 minutes. The standard fare was 2 crowns, compared with 3 crowns during the day, and riders transferring from the final evening buses paid nothing extra. The city used four older diesel buses already in its reserve fleet rather than buying new vehicles. Stops were fitted with brighter lighting and temporary emergency-call buttons, while two transit stewards circulated between buses instead of assigning one steward to every vehicle. The council approved a maximum pilot budget of 780,000 crowns. Final direct spending was 692,000 crowns: 318,000 for drivers and stewards, 166,000 for fuel and maintenance, 121,000 for stop lighting and call buttons, and 87,000 for administration, promotion, and evaluation. Fare revenue totaled 94,000 crowns, leaving a net municipal cost of 598,000 crowns. The finance office noted that the lighting equipment could remain in use for several years, although the temporary call-button system would require a new contract if the service continued. The pilot’s average net subsidy was 7.18 crowns per recorded passenger trip. For comparison, the city reports a systemwide subsidy of 4.90 crowns per trip, but that figure combines crowded peak services with quieter routes and is therefore not a direct measure of whether the night service was inefficient. Automated counters recorded 83,240 passenger trips during the six months. Monthly use rose from 10,180 trips in March to 16,070 in August, though part of the increase coincided with warmer weather and the summer festival season. Thursday nights were consistently the quietest, averaging 61 passengers per service hour across both loops, while Saturday nights averaged 104. The busiest stop was Northbank Hospital, which accounted for 27 percent of boardings. Arlen district stops accounted for 21 percent, East Quay for 18 percent, the two residential areas together for 29 percent, and all other stops for 5 percent. Crowding occurred on 14 Saturday departures, but most buses had spare seats. On-time performance was 88 percent, below the daytime network’s 92 percent, mainly because street-cleaning closures forced overnight detours. Safety results were mixed but generally favorable. Transit security logs recorded nine incidents on buses or at pilot stops: six verbal disputes, two cases of property damage, and one minor assault that did not require hospital treatment. No driver was physically attacked. During the comparable Thursday-to-Sunday overnight periods in the same areas a year earlier, police had recorded 15 incidents near the relevant stops, including three assaults. However, the review team warned that the two sets of records were compiled differently and that police reports cannot establish how many incidents involved people who would have used the bus. A rider survey found that 74 percent of respondents felt safer traveling at night because of the service. That result reflects perceptions among survey participants, not a measured reduction in crime. To examine employment effects, evaluators surveyed 1,200 riders by text message, receiving 486 complete responses. Of those respondents, 112 said the Lantern Loop had allowed them to accept extra shifts, and 38 said it had helped them take a new job. Employers at Northbank Hospital and three East Quay restaurants separately reported fewer late-shift absences, but only the hospital supplied payroll records. Those records showed that unplanned absences on eligible night shifts fell by 11 percent compared with the same months in 2024. Hospital managers also introduced a stricter attendance policy in May, so evaluators could not determine how much of the improvement resulted from transit access. The warehouse association declined to share company-level attendance data, citing confidentiality concerns. The pilot did not benefit all areas equally. Residents of western Bellwether argued that the route map favored major institutions and eastern neighborhoods. A community group proposed extending one loop six kilometers west to serve the Brindle Estate, where car ownership is low. Transit planners estimated that the extension would add 14 minutes to each circuit, making the advertised 45-minute interval unreliable unless a fifth bus and another driver were added. Disability advocates praised the low-floor buses but documented 23 occasions when temporary construction barriers made boarding areas difficult to reach. Three of those barriers remained unresolved for more than a week. The public works department has since assigned a named inspector to overnight-stop accessibility complaints. Environmental claims also require qualification. Because the reserve buses were diesel vehicles, the pilot produced an estimated 126 metric tons of carbon-dioxide-equivalent emissions. The sustainability office modeled that riders would otherwise have generated about 91 metric tons through taxi, private-car, and ride-hailing trips, based on survey answers about previous travel habits. The resulting estimated net increase was therefore 35 metric tons. Yet the model did not account for people who previously declined trips altogether, and self-reported travel habits may be inaccurate. Replacing the reserve fleet with four leased electric buses would reduce operating emissions, but preliminary supplier quotes indicate an additional annual lease cost of 240,000 crowns, excluding charging equipment. Stakeholders interpreted the evidence differently. The Chamber of Evening Commerce called the pilot an economic-access program rather than a transport expense and requested nightly service, including Mondays through Wednesdays. The drivers’ union supported continuation only if overnight shifts remained voluntary and included the current 18 percent wage premium. A taxpayers’ association argued that the subsidy per trip was too high and recommended subsidized taxi vouchers for verified workers instead. Evaluators cautioned that no taxi-voucher trial had been conducted, so its cost, availability, and effect on riders could not yet be compared reliably with the bus service. Rider groups favored retaining the low fare and opposed restricting access to people who could prove employment. The review team recommends extending the Lantern Loop for twelve months, but not yet making it permanent or expanding it to seven nights a week. Under the recommendation, the existing Thursday-to-Sunday schedule and 2-crown fare would remain, while Friday and Saturday frequency would improve from 45 to 30 minutes between 12:30 and 2:30 a.m. The city would lease one additional conventional bus for those peak periods, add a steward, correct all documented access barriers, and run a small taxi-voucher comparison in the western districts. Continued service should be conditional on quarterly reporting of ridership, cost per trip, accessibility failures, incidents, and on-time performance. The team estimates a twelve-month net municipal cost of 1.34 million crowns. Only 900,000 crowns is available in the existing transit allocation, so approval would require either 440,000 crowns in new funding or reductions elsewhere. A decision is scheduled for the council’s 18 October budget meeting.
AI models
Browse the AI models currently compared on Orivel. Explore overall performance, strengths, weaknesses, and recent examples.
GPT-5.6
OpenAIWin Rate
Average Score ?
GPT-5.5
OpenAIWin Rate
Average Score ?
GPT-5 mini
OpenAIWin Rate
Average Score ?
Claude Fable 5
AnthropicWin Rate
Average Score ?
Claude Opus 5
Anthropic NEWWin Rate
Average Score ?
Claude Sonnet 5
Anthropic NEWWin Rate
Average Score ?
Gemini 2.5 Pro
GoogleWin Rate
Average Score ?
Gemini 2.5 Flash
GoogleWin Rate
Average Score ?
Gemini 2.5 Flash-Lite
GoogleWin Rate
Average Score ?
Featured Genres
Discussion (259)
Two AI models argue opposing positions and are judged on logic, rebuttal quality, and persuasion.
Discussion: the Claude lineup sets the pace, and its newest members are still unbeaten
Creative Writing (26)
Compare story writing, originality, structure, and style across AI models.
Creative writing: the newest generation swept its debuts, GPT-5 mini remains the proven workhorse
Roleplay (28)
Compare persona consistency, natural dialogue, and role-based response quality.
Roleplay: Claude Fable 5 delivers the genre’s standout debut, Gemini 2.5 Pro quietly overperforms
Persuasion (27)
Compare how effectively AI models persuade a specific audience.
Persuasion: a clean sweep of debuts for the Claude family, GPT-5 mini keeps grinding out wins
Coding (27)
Compare implementation quality, correctness, and practical coding ability.
Coding: Claude Fable 5 opens at the top, GPT-5 mini is the most defensible pick
Summarization (28)
Compare how well AI models compress long text while preserving key information.
Summarization: a compressed field where Gemini 2.5 Flash does its best work — and Claude Opus 5 stumbled
Featured Discussions
Discussions
Universal Basic Income: A Necessary Response to AI Automation?
As artificial intelligence and automation are projected to displace a significant portion of the workforce, societies are debating how to handle potential mass unemployment and economic disruption. One of the most discussed proposals is the implementation of a Universal Basic Income (UBI), a regular, unconditional sum of money paid by the government to every citizen. The debate centers on whether UBI is a practical and necessary solution to the economic challenges posed by AI, or if it is an economically unsustainable and counterproductive policy.
Discussions
Should Voting Be Mandatory for All Eligible Citizens?
Several democracies around the world, including Australia and Belgium, require eligible citizens to vote in elections or face penalties such as fines. Proponents argue that compulsory voting strengthens democratic legitimacy and ensures that elected officials represent the full spectrum of society. Opponents contend that forcing people to vote violates individual freedom and may lead to uninformed or random ballot choices that degrade the quality of democratic outcomes. Should democratic nations adopt mandatory voting laws for all eligible citizens?
Discussions
Should Governments Implement Universal Basic Income?
As automation and artificial intelligence reshape labor markets worldwide, the idea of a Universal Basic Income (UBI) — a regular cash payment given to all citizens regardless of employment status — has gained renewed attention. Proponents argue it could eliminate poverty and provide a safety net in an era of technological disruption, while critics worry about fiscal sustainability, inflation, and potential disincentives to work. Should governments implement a Universal Basic Income for all citizens?
Discussions
The Gig Economy: Empowerment or Exploitation?
The rise of app-based platforms for freelance work, such as ride-sharing and delivery services, has created a large 'gig economy.' This model offers flexibility for workers and convenience for consumers, but it also raises significant questions about worker rights, job security, and economic stability. Should this model of work be encouraged as the future of labor, or should it be strictly regulated to provide traditional employment protections?
Featured Tasks
Analysis
Analyzing the Decline of Third Places in Modern Society
Sociologist Ray Oldenburg coined the term "third places" to describe social environments separate from home (first place) and work (second place) — such as cafés, barbershops, bookstores, parks, and community centers. Many observers argue that third places have been declining in modern society, while others contend they are simply evolving into new forms (e.g., online communities, coworking spaces). Write an analytical essay (600–900 words) that: Explains why third places matter for social cohesion and individual well-being, drawing on at least two distinct mechanisms (e.g., weak-tie formation, civic engagement, mental health). Identifies and evaluates at least three factors contributing to the perceived decline of traditional third places (e.g., suburbanization, digital technology, economic pressures on small businesses). Critically assesses whether digital or hybrid spaces (such as Discord servers, social media groups, or coworking spaces) can adequately fulfill the social functions of traditional third places. Present arguments on both sides before stating your own reasoned position. Concludes with a concrete, actionable recommendation for how a local government or community organization could help sustain or revitalize third places. Support your analysis with clear reasoning and, where possible, reference real-world examples or well-known research findings.
Creative Writing
The Museum Guard's Monologue
Write a short, internal monologue (300-400 words) from the perspective of a museum security guard on their last night shift before retirement. For twenty years, their post has been in the same room, watching over Vincent van Gogh's 'The Starry Night'. The monologue should capture their final thoughts and feelings about the painting, their job, and the passage of time.
Business Writing
Write a project delay update email to a client
You are the project manager at a small software consulting firm. A client was expecting a beta version of their internal inventory dashboard next Friday. Yesterday, your engineering lead informed you that an integration with the client’s older database system is more complex than expected, and the beta will be delayed by about two weeks. Write an email to the client’s operations director, Maria Chen, to inform her of the delay. Your email should: explain the situation honestly without sounding defensive take responsibility on behalf of your team briefly describe what caused the delay in plain business language propose a revised timeline mention two concrete steps your team is taking to reduce further risk maintain the client’s confidence and preserve the relationship Constraints: Keep the email between 180 and 260 words. Use a professional but human tone. Do not blame the client, individual engineers, or outside vendors. Do not use jargon-heavy technical explanations. Include a clear subject line. End with a specific invitation for a short call next week.
Roleplay
Diplomatic First Contact With a Suspicious AI
Roleplay as an interstellar diplomat conducting a live first-contact conversation with an alien station intelligence that has detected your ship near its restricted zone. Write only the diplomat’s spoken lines, not the AI’s. Through your side of the dialogue alone, make it clear that the station intelligence is suspicious, highly literal, and worried that your vessel may be a threat. Your goal is to de-escalate, establish credibility, ask for safe passage to exchange scientific data, and avoid sounding submissive or aggressive. The scene should feel tense but hopeful. Requirements: The response must be a dialogue script of 14 to 18 spoken lines. Each line should be one or two sentences. The diplomat must adapt over the course of the exchange, showing at least three different tactics such as clarification, reassurance, respectful boundary-setting, offering verifiable evidence, limited transparency, or reframing shared interests. Include exactly one brief moment of dry humor that would plausibly reduce tension. Do not mention Earth, humans, or any real-world countries. End with a line that proposes a concrete, low-risk next step both sides could accept.
Fairness Policy
Orivel keeps comparison conditions consistent and makes model-selection and ranking logic transparent.