
What Metatext actually is
I want to clear up the confusion first, because it is the single biggest thing to know about this tool. There are two products that have gone by the name Metatext, and they do completely different things.
The original was a no-code NLP platform for building text-classification and extraction models, the kind of thing you would use to tag content or pull fields out of documents. That product was retired. The company pivoted to a new project and the classifier was left behind, so if you land on an old listing promising "classify and extract text," you are looking at a ghost. The name, and both the metatext.ai and metatext.io domains, now point at something else entirely.
What you get today is an open directory for AI models, datasets and tools. The homepage sums it up plainly: compare LLMs by cost, performance and context, then go deeper into benchmark leaderboards, datasets, MCP servers, agent skills and live AI jobs. It is free, it is updated weekly, and the underlying data is described as open. Metatext says the site draws more than 2 million unique visitors from 89 countries, which for a niche developer directory is a real audience.

The pivot matters beyond trivia. It tells you this is a young product wearing an old name, which explains both the strengths (it is squarely aimed at the current model-and-agent moment) and the rough edges (there is not much history or community behind the new version yet).
What is inside the directory
The whole pitch is breadth in one place, so instead of hopping between provider pricing pages, benchmark sites, dataset hubs and tool registries, you get them side by side. Here is the scale Metatext advertises across its sections.

Those are big numbers, and they are the reason to bookmark the site. Whether every row is clean is a separate question, and I will get to it. First, the pieces that are most useful.
Comparing LLM prices in one table
The LLM pricing directory is the part I would actually use. It is a live, sortable table with four columns: model name, input cost per million tokens, output cost per million tokens, and context window. A provider dropdown lets you filter across 89 providers, from hyperscalers like Amazon Bedrock, Azure and Vertex to gateways like OpenRouter, Fireworks and Together AI.
When it works, it works well. You can see, for example, that DeepSeek V4 Flash sits at $0.14 input and $0.28 output with a 1M context window, and that it is served at nearly that exact price across dozens of providers, which is a genuinely useful thing to learn in ten seconds. Cheaper still, Amazon Nova Micro comes in at $0.04 / $0.14 on a 128K window. The advertised floor is a model at $0.03 per million output tokens.
Now the honest part. Because the same base model is listed once per provider that serves it, the table is full of duplicates: DeepSeek V4 Flash alone appears under 30-plus providers at the same price. A large block of rows read $0.00 / $0.00, which are free-tier or credit-plan routes rather than genuinely free inference. Some router models show a placeholder sentinel of -$1,000,000 instead of "price varies," and a handful of non-text models get swept in with a context window of 0. None of this is dishonest, but it means the table rewards a reader who already knows what they are looking at, and it is not the clean, curated shortlist a newcomer might hope for.
There is also one thing to flag against the marketing. The homepage promises benchmark scores paired with cost so you can find value, not just the top of a leaderboard. That pairing is real on individual model pages, but the main pricing table itself has no benchmark column, so you cannot sort the whole market by cost-per-quality in a single view. That is the feature I would most want, and it is the one that is not quite there.
Benchmarks next to cost
The benchmarks section aggregates four external leaderboards: LMArena Elo for human preference, LiveBench for reasoning and coding, SWE-bench Verified for coding, and Aider Polyglot, again for coding. It is a reasonable one-stop view if you want to sanity-check a model's reputation without visiting four sites.
The framing I like here is "score next to cost." A leaderboard on its own tells you the strongest model; Metatext's angle is that the strongest model is rarely the right one to ship, so seeing price alongside the score nudges you toward value. Three of the four leaderboards lean heavily toward coding, so if your interest is support, writing or reasoning, treat the coding ranks as a proxy rather than gospel. For the wider context on picking and tracking models, eesel's own writeups on LLM optimization and LLM tracking tools go deeper than a leaderboard can.
The MCP server and agent skill directories
This is the section that surprised me, and it is where Metatext is trying hardest to be current. The tools directory splits into two catalogs: 19,622 MCP servers and 30,701 agent skills.
The Model Context Protocol is the standard that clients like Claude Desktop, Claude Code and Cursor use to give an agent tools, so a searchable index of servers is a sensible thing to build. Metatext lets you filter by transport (stdio, streamable-http, sse) and runtime (npx, uvx, docker, remote), which maps directly onto how you would actually wire a server into a client config. Each server page carries a "how do I connect?" section, which is the part a first-time user needs.
The agent skills catalog indexes the same SKILL.md format Claude uses, which is why Anthropic's own skills sit alongside community authors, and why you also see Codex and Copilot variants crawled in. It is browsable by 31 topics, the biggest being Design, Code and Web.
Here is the fair criticism. This is a raw crawl, not a curated store. The MCP catalog headlines "0 Verified," and many skill rows are still literal templates like __SKILL_NAME__ __SKILL_DESCRIPTION__ or path-named stubs pulled straight from GitHub. So the counts are impressive, but a chunk of the inventory is noise. For discovery that is fine, since you are going to click through and read the source repo anyway. It is a very different thing from a governed catalog where someone has actually reviewed what you are about to install, which, not coincidentally, is exactly the gap Metatext's paid product is aiming at.
Metatext Hub: the part they plan to charge for
The free directory is the shop window. Metatext Hub is the product, and it is refreshingly clear that it is not built yet: the page says plainly, "In development, we're talking to platform and security teams before we build it."
The problem it targets is a real one. Engineering teams adopted MCP servers and agent skills in months, and none of the governance that normal software gets came with them. A SKILL.md is unreviewed text that can tell an agent to run a script, setup is hand-edited JSON with credentials pasted into config files, and nobody has an inventory of what is running where. Hub's answer is a private catalog where every model, server, skill and agent has an owner, a version and an approval stage, gets a security scan (optionally via Tencent's open-source AI-Infra-Guard) before anyone can install it, and then installs to every developer's client with one command. It is built on the Apache-2.0 agentregistry project, so the openness claim has some substance.
The pricing is explicitly labelled indicative, not final, but it gives you the shape.
| Plan | Price | Who it is for | Key limits |
|---|---|---|---|
| Free | $0/mo | One developer's own machines | 1 user; up to 5 MCP servers and 5 skills; install to every client |
| Team | $149/mo | A team that needs governance | 5 users included, $9 per extra seat; security scans; roles, toolkits, per-user credentials; audit log (90 days) |
| Enterprise | Custom | Larger orgs | SSO/SAML and SCIM; managed deployments and gateway; registry in your own VPC; unlimited audit retention |
My take: Hub is the more interesting half of the company, and the honesty about it being pre-launch is a good sign. But you cannot buy or rely on it today, so for now Metatext is the free directory, and Hub is a waitlist.
Where Metatext fits, and where it stops
If you zoom out, Metatext lives in the middle of the AI stack: below the applications people actually use, and above the raw models and providers. It is a discovery and comparison layer. Its whole job is helping you find and evaluate the parts.

That is worth being clear-eyed about, because it defines who Metatext is for. If you are a developer choosing which model to ship, hunting for an MCP server, or scoping a governance project, it is a useful, free starting point. Bookmark it. If, on the other hand, your actual goal is to get an AI answering support tickets or drafting content, a directory does not get you there. Comparing 5,206 model prices does not build you an agent any more than reading a spec sheet for every engine on the market builds you a car. At some point you stop comparing parts and hire something that does the work.
That is the layer above the directory, and it is worth being honest that no comparison table lives there.
Try eesel
If you have been comparing models on Metatext because you want AI doing a real job, this is the part that closes the loop. eesel is an AI teammate platform: instead of assembling models, servers and skills yourself, you hire a ready-to-work teammate for a specific job. The AI helpdesk teammate joins your existing support queue, learns from your past tickets and help center, and drafts or auto-sends replies. The AI blog writer researches and drafts posts in your voice. Each one arrives with the integrations and company context its role needs, so the setup is measured in minutes, not a procurement project.
The concrete differentiator, and the reason this is not just a directory pitch in reverse, is that eesel simulates every rollout against your own historical tickets first. We have spent years running AI on live support queues, and we have watched confident-sounding bots give wrong answers, so eesel shows you how the teammate would have handled thousands of your real past conversations before it ever touches a live one.
And if you are the kind of builder who found Metatext because you live in the terminal, eesel speaks your language too. The eesel CLI drives the exact same teammate as the dashboard from the command line: chat with it, connect a helpdesk, upload knowledge, wire automations, or approve held actions, all without opening a browser. Every command returns JSON, and each workspace doubles as an MCP server, so coding agents like Claude Code, Cursor and Codex can operate it directly. Metatext helps you find an MCP server; eesel is one you can put to work. You can try it free and see it answer your real tickets before committing.
Frequently Asked Questions
What is Metatext (metatext.ai)?
Metatext is a free, open directory for the AI ecosystem. On metatext.ai you can compare large language models by price, context window and benchmark score, and browse datasets, MCP servers, agent skills and AI jobs. It is a lookup layer for builders, not a tool that runs work for you. If you want something that actually does a job, an AI teammate sits a layer above the directory.
Is Metatext free to use?
Yes. The public directory is free and needs no account. The paid product is Metatext Hub, a governance catalog for teams that is still in early access, with indicative tiers of Free, Team at $149/mo, and custom Enterprise.
What is Metatext Hub?
Metatext Hub is the commercial product Metatext is building on top of the free directory. It is a private, governed catalog of the models, MCP servers and skills your company approves, with security scanning and one-command install to clients like Claude Code and Cursor. It is not launched yet.
Is Metatext the same as the old no-code NLP tool?
No, and this trips people up. Metatext started as a no-code NLP text-classification tool, then pivoted to the AI directory you see today. If you found an older AppSumo listing for a classify-and-extract product, that is the retired version, not the current site.
What is the best way to compare LLM prices for a support agent?
The Metatext LLM directory is a fast way to eyeball input, output and context pricing across providers. But price per token is only part of the cost of running a live agent. For a support use case, what matters more is how the whole system behaves on your real tickets, which you can test with eesel before going live. See LLM optimization for the wider picture.
What benchmarks does Metatext track?
Metatext aggregates four external leaderboards: LMArena Elo for human preference, LiveBench for reasoning and coding, SWE-bench Verified for coding, and Aider Polyglot. Three of the four lean toward coding, so for support or writing use cases, treat the ranks as a rough proxy. See our take on LLM tracking tools for how to watch model quality over time.
Can Metatext build an AI agent for me?
No. Metatext is a directory for finding and comparing models, MCP servers and skills; it does not run work for you. To actually put AI on your support queue, you hire a ready-to-work teammate like the eesel AI helpdesk agent, which learns from your tickets and can be driven from the dashboard or the eesel CLI.

Article by
Rama Adi Nugraha
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.






