Metatext review: inside metatext.ai, the open AI model directory

Rama Adi Nugraha
Written by

Rama Adi Nugraha

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 28, 2026

Expert Verified
Illustration of a catalog dashboard comparing AI models, datasets and tools side by side

What Metatext actually is

I want to clear up the confusion first, because it is the single biggest thing to know about this tool. There are two products that have gone by the name Metatext, and they do completely different things.

The original was a no-code NLP platform for building text-classification and extraction models, the kind of thing you would use to tag content or pull fields out of documents. That product was retired. The company pivoted to a new project and the classifier was left behind, so if you land on an old listing promising "classify and extract text," you are looking at a ghost. The name, and both the metatext.ai and metatext.io domains, now point at something else entirely.

What you get today is an open directory for AI models, datasets and tools. The homepage sums it up plainly: compare LLMs by cost, performance and context, then go deeper into benchmark leaderboards, datasets, MCP servers, agent skills and live AI jobs. It is free, it is updated weekly, and the underlying data is described as open. Metatext says the site draws more than 2 million unique visitors from 89 countries, which for a niche developer directory is a real audience.

Before and after: Metatext pivoted from a retired no-code NLP text classifier to an open directory covering LLM pricing, benchmarks, datasets, MCP servers, agent skills and AI jobs
Before and after: Metatext pivoted from a retired no-code NLP text classifier to an open directory covering LLM pricing, benchmarks, datasets, MCP servers, agent skills and AI jobs

The pivot matters beyond trivia. It tells you this is a young product wearing an old name, which explains both the strengths (it is squarely aimed at the current model-and-agent moment) and the rough edges (there is not much history or community behind the new version yet).

What is inside the directory

The whole pitch is breadth in one place, so instead of hopping between provider pricing pages, benchmark sites, dataset hubs and tool registries, you get them side by side. Here is the scale Metatext advertises across its sections.

The six sections of the Metatext directory with their catalog counts: 5,206 priced LLMs, 7,967 models, 5,413 datasets, 19,622 MCP servers, 30,701 agent skills, and 244 AI jobs
The six sections of the Metatext directory with their catalog counts: 5,206 priced LLMs, 7,967 models, 5,413 datasets, 19,622 MCP servers, 30,701 agent skills, and 244 AI jobs

Those are big numbers, and they are the reason to bookmark the site. Whether every row is clean is a separate question, and I will get to it. First, the pieces that are most useful.

Comparing LLM prices in one table

The LLM pricing directory is the part I would actually use. It is a live, sortable table with four columns: model name, input cost per million tokens, output cost per million tokens, and context window. A provider dropdown lets you filter across 89 providers, from hyperscalers like Amazon Bedrock, Azure and Vertex to gateways like OpenRouter, Fireworks and Together AI.

The Metatext LLM pricing table on metatext.ai, comparing input cost, output cost and context window across providers

When it works, it works well. You can see, for example, that DeepSeek V4 Flash sits at $0.14 input and $0.28 output with a 1M context window, and that it is served at nearly that exact price across dozens of providers, which is a genuinely useful thing to learn in ten seconds. Cheaper still, Amazon Nova Micro comes in at $0.04 / $0.14 on a 128K window. The advertised floor is a model at $0.03 per million output tokens.

Now the honest part. Because the same base model is listed once per provider that serves it, the table is full of duplicates: DeepSeek V4 Flash alone appears under 30-plus providers at the same price. A large block of rows read $0.00 / $0.00, which are free-tier or credit-plan routes rather than genuinely free inference. Some router models show a placeholder sentinel of -$1,000,000 instead of "price varies," and a handful of non-text models get swept in with a context window of 0. None of this is dishonest, but it means the table rewards a reader who already knows what they are looking at, and it is not the clean, curated shortlist a newcomer might hope for.

There is also one thing to flag against the marketing. The homepage promises benchmark scores paired with cost so you can find value, not just the top of a leaderboard. That pairing is real on individual model pages, but the main pricing table itself has no benchmark column, so you cannot sort the whole market by cost-per-quality in a single view. That is the feature I would most want, and it is the one that is not quite there.

Benchmarks next to cost

The benchmarks section aggregates four external leaderboards: LMArena Elo for human preference, LiveBench for reasoning and coding, SWE-bench Verified for coding, and Aider Polyglot, again for coding. It is a reasonable one-stop view if you want to sanity-check a model's reputation without visiting four sites.

The Metatext benchmarks page aggregating LMArena, LiveBench, SWE-bench and Aider leaderboards

The framing I like here is "score next to cost." A leaderboard on its own tells you the strongest model; Metatext's angle is that the strongest model is rarely the right one to ship, so seeing price alongside the score nudges you toward value. Three of the four leaderboards lean heavily toward coding, so if your interest is support, writing or reasoning, treat the coding ranks as a proxy rather than gospel. For the wider context on picking and tracking models, eesel's own writeups on LLM optimization and LLM tracking tools go deeper than a leaderboard can.

The MCP server and agent skill directories

This is the section that surprised me, and it is where Metatext is trying hardest to be current. The tools directory splits into two catalogs: 19,622 MCP servers and 30,701 agent skills.

The Metatext MCP servers directory on metatext.ai, filterable by transport, runtime and author

The Model Context Protocol is the standard that clients like Claude Desktop, Claude Code and Cursor use to give an agent tools, so a searchable index of servers is a sensible thing to build. Metatext lets you filter by transport (stdio, streamable-http, sse) and runtime (npx, uvx, docker, remote), which maps directly onto how you would actually wire a server into a client config. Each server page carries a "how do I connect?" section, which is the part a first-time user needs.

The agent skills catalog indexes the same SKILL.md format Claude uses, which is why Anthropic's own skills sit alongside community authors, and why you also see Codex and Copilot variants crawled in. It is browsable by 31 topics, the biggest being Design, Code and Web.

Here is the fair criticism. This is a raw crawl, not a curated store. The MCP catalog headlines "0 Verified," and many skill rows are still literal templates like __SKILL_NAME__ __SKILL_DESCRIPTION__ or path-named stubs pulled straight from GitHub. So the counts are impressive, but a chunk of the inventory is noise. For discovery that is fine, since you are going to click through and read the source repo anyway. It is a very different thing from a governed catalog where someone has actually reviewed what you are about to install, which, not coincidentally, is exactly the gap Metatext's paid product is aiming at.

Metatext Hub: the part they plan to charge for

The free directory is the shop window. Metatext Hub is the product, and it is refreshingly clear that it is not built yet: the page says plainly, "In development, we're talking to platform and security teams before we build it."

The Metatext Hub early-access page describing a governed catalog for a company's AI tools

The problem it targets is a real one. Engineering teams adopted MCP servers and agent skills in months, and none of the governance that normal software gets came with them. A SKILL.md is unreviewed text that can tell an agent to run a script, setup is hand-edited JSON with credentials pasted into config files, and nobody has an inventory of what is running where. Hub's answer is a private catalog where every model, server, skill and agent has an owner, a version and an approval stage, gets a security scan (optionally via Tencent's open-source AI-Infra-Guard) before anyone can install it, and then installs to every developer's client with one command. It is built on the Apache-2.0 agentregistry project, so the openness claim has some substance.

The pricing is explicitly labelled indicative, not final, but it gives you the shape.

PlanPriceWho it is forKey limits
Free$0/moOne developer's own machines1 user; up to 5 MCP servers and 5 skills; install to every client
Team$149/moA team that needs governance5 users included, $9 per extra seat; security scans; roles, toolkits, per-user credentials; audit log (90 days)
EnterpriseCustomLarger orgsSSO/SAML and SCIM; managed deployments and gateway; registry in your own VPC; unlimited audit retention

My take: Hub is the more interesting half of the company, and the honesty about it being pre-launch is a good sign. But you cannot buy or rely on it today, so for now Metatext is the free directory, and Hub is a waitlist.

Where Metatext fits, and where it stops

If you zoom out, Metatext lives in the middle of the AI stack: below the applications people actually use, and above the raw models and providers. It is a discovery and comparison layer. Its whole job is helping you find and evaluate the parts.

A three-layer diagram: the model and provider layer at the bottom, a highlighted discovery layer of directories in the middle, and the application layer of agents that do the work on top
A three-layer diagram: the model and provider layer at the bottom, a highlighted discovery layer of directories in the middle, and the application layer of agents that do the work on top

That is worth being clear-eyed about, because it defines who Metatext is for. If you are a developer choosing which model to ship, hunting for an MCP server, or scoping a governance project, it is a useful, free starting point. Bookmark it. If, on the other hand, your actual goal is to get an AI answering support tickets or drafting content, a directory does not get you there. Comparing 5,206 model prices does not build you an agent any more than reading a spec sheet for every engine on the market builds you a car. At some point you stop comparing parts and hire something that does the work.

That is the layer above the directory, and it is worth being honest that no comparison table lives there.

Try eesel

If you have been comparing models on Metatext because you want AI doing a real job, this is the part that closes the loop. eesel is an AI teammate platform: instead of assembling models, servers and skills yourself, you hire a ready-to-work teammate for a specific job. The AI helpdesk teammate joins your existing support queue, learns from your past tickets and help center, and drafts or auto-sends replies. The AI blog writer researches and drafts posts in your voice. Each one arrives with the integrations and company context its role needs, so the setup is measured in minutes, not a procurement project.

The eesel AI helpdesk agent page, showing an AI teammate that plugs into your existing support tools

The concrete differentiator, and the reason this is not just a directory pitch in reverse, is that eesel simulates every rollout against your own historical tickets first. We have spent years running AI on live support queues, and we have watched confident-sounding bots give wrong answers, so eesel shows you how the teammate would have handled thousands of your real past conversations before it ever touches a live one.

And if you are the kind of builder who found Metatext because you live in the terminal, eesel speaks your language too. The eesel CLI drives the exact same teammate as the dashboard from the command line: chat with it, connect a helpdesk, upload knowledge, wire automations, or approve held actions, all without opening a browser. Every command returns JSON, and each workspace doubles as an MCP server, so coding agents like Claude Code, Cursor and Codex can operate it directly. Metatext helps you find an MCP server; eesel is one you can put to work. You can try it free and see it answer your real tickets before committing.

Frequently Asked Questions

What is Metatext (metatext.ai)?

Metatext is a free, open directory for the AI ecosystem. On metatext.ai you can compare large language models by price, context window and benchmark score, and browse datasets, MCP servers, agent skills and AI jobs. It is a lookup layer for builders, not a tool that runs work for you. If you want something that actually does a job, an AI teammate sits a layer above the directory.

Is Metatext free to use?

Yes. The public directory is free and needs no account. The paid product is Metatext Hub, a governance catalog for teams that is still in early access, with indicative tiers of Free, Team at $149/mo, and custom Enterprise.

What is Metatext Hub?

Metatext Hub is the commercial product Metatext is building on top of the free directory. It is a private, governed catalog of the models, MCP servers and skills your company approves, with security scanning and one-command install to clients like Claude Code and Cursor. It is not launched yet.

Is Metatext the same as the old no-code NLP tool?

No, and this trips people up. Metatext started as a no-code NLP text-classification tool, then pivoted to the AI directory you see today. If you found an older AppSumo listing for a classify-and-extract product, that is the retired version, not the current site.

What is the best way to compare LLM prices for a support agent?

The Metatext LLM directory is a fast way to eyeball input, output and context pricing across providers. But price per token is only part of the cost of running a live agent. For a support use case, what matters more is how the whole system behaves on your real tickets, which you can test with eesel before going live. See LLM optimization for the wider picture.

What benchmarks does Metatext track?

Metatext aggregates four external leaderboards: LMArena Elo for human preference, LiveBench for reasoning and coding, SWE-bench Verified for coding, and Aider Polyglot. Three of the four lean toward coding, so for support or writing use cases, treat the ranks as a rough proxy. See our take on LLM tracking tools for how to watch model quality over time.

Can Metatext build an AI agent for me?

No. Metatext is a directory for finding and comparing models, MCP servers and skills; it does not run work for you. To actually put AI on your support queue, you hire a ready-to-work teammate like the eesel AI helpdesk agent, which learns from your tickets and can be driven from the dashboard or the eesel CLI.

Share this article

Rama Adi Nugraha

Article by

Rama Adi Nugraha

Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.

Related Posts

All posts →
Illustration of a pricing page with three plan cards next to an AI model directory dashboard
Trending

Metatext pricing 2026: what the AI directory actually costs

A clear breakdown of Metatext pricing: the metatext.ai directory is free, the paid tiers on the pricing page are template placeholders, and the real paid product, Metatext Hub, is still in early access.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 28, 2026
Banner image for Claude Sonnet 4.6 review: The sweet spot between performance and price
Guides

Claude Sonnet 4.6 review: The sweet spot between performance and price

Anthropic's Claude Sonnet 4.6 punches above its weight class with frontier-level coding performance, a 1M token context window, and significant improvements over Sonnet 4.5.

Stevia PutriStevia PutriFeb 26, 2026
Editorial illustration of a decision engine being scored in review, in Laya terracotta
Trending

Laya AI review: the open 33ms decision model, tested honestly

An honest Laya AI review: the open-source 421M decision model is genuinely fast and calibrated, but the headline accuracy hides a fine-tuning caveat. Here is where it wins, where it does not, and who should use it.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 22, 2026
Editorial illustration of a fast decision engine sorting typed answers, in Laya terracotta
Trending

Laya AI: the open 33ms decision model that can't hallucinate

Laya AI is Convai's open-source, 421M System 1 decision model: 33ms typed answers with calibrated confidence, no text generation, and an honest zero-shot caveat.

Alicia Kirana UtomoAlicia Kirana UtomoSep 22, 2026
MiniCPM5-2B, a compact 2B open-weight model that runs on phones and laptops
Trending

MiniCPM5-2B: a 2B open model that runs on-device and beats bigger ones

A close look at MiniCPM5-2B: what OpenBMB's compact 2B model actually is, how it scores, where it runs, and what a raw open model still needs to do real work.

Alicia Kirana UtomoAlicia Kirana UtomoSep 9, 2026
Illustration of Inkling, Thinking Machines Lab's open-weights AI model under review
Trending

Inkling review: is Thinking Machines' open model worth it?

An honest Inkling review: what Thinking Machines Lab's first open-weights model is genuinely good at, where the price and benchmarks let it down, and who should actually run it.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 20, 2026
Illustration of Inkling, Thinking Machines Lab's open-weights AI model
Trending

Inkling explained: Thinking Machines' open-weights AI model

What Inkling actually is: Thinking Machines Lab's first open-weights model, its real benchmarks, what it costs to run, and whether it belongs anywhere near a support queue.

Alicia Kirana UtomoAlicia Kirana UtomoJul 20, 2026
Gemini vs Claude: Which AI model is right for you in 2025?
Guides

Gemini vs Claude: Which AI model is right for you in 2026?

Gemini vs Claude: explore the strengths, differences, and key features of each AI to discover which one best fits your needs.

Stevia PutriStevia PutriAug 17, 2025
Grok 4.7 review hero banner with the Grok logo on a dark abstract compute backdrop
Trending

Grok 4.7 review: is xAI's cheap frontier model actually good?

A hands-on Grok 4.7 review: what it's good at, where it still loses to Fable 5.1 and GPT-5.6 Sol, the real pricing, the Fast-variant tax, and who should actually run it.

Rama Adi NugrahaRama Adi NugrahaSep 23, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free