Fantasy Agent Get started

How it works

No black box. Here's every input, every judgement, and every check.

Does all the research with your brain, so you don't have to.

Most fantasy tools hand you the same consensus ranking as everyone else in your league. This one reads your league, reasons about each player's whole situation, follows the philosophy you wrote, and then grades itself against what actually happened.

How it connects to your league

Setup is three steps and takes a few minutes.

  1. Verify the league is yours. You add a short code to your team name in that league for a minute. We read it back from Sleeper's public API to confirm ownership, then you remove it.
  2. Write your philosophy. A short questionnaire turns into the document the agent argues from: floor over ceiling, how aggressive to be with FAAB, whether you chase matchups. Each league can override it.
  3. Pick your days and time. Tuesday sends the full report; Thursday, Saturday and Sunday send lineup checks. Everything arrives at the hour you choose, in your own time zone.

We never touch your roster. Sleeper's API is read-only, so no password, token or league invite is ever needed, and the agent could not submit a lineup or claim even if it wanted to. Every recommendation ends with you making the move yourself.

ESPN leagues can be connected today, but the agent's data adapter is Sleeper-only for now, so reports run for Sleeper leagues.

Where the data comes from

Every input is a public source, read directly. There is no resold data feed in the middle.

SourceWhat it providesRefresh
Your leagueRosters, starters and bench, lineup slots, league scoring and roster rules, FAAB budgets, transactions, your weekly opponent, injury designations with body part and notes, depth-chart order, trending adds, and weekly usage stats including team totalsEvery run
Play-by-play and tracking dataThe NFL schedule, offensive-line snap counts, and player tracking stats such as separation, air-yard share, rushing yards over expected and time to throwCached about a day
Official weather forecastsKickoff conditions for uncovered stadiumsPer run, inside the shared player brief
Betting marketsGame totals and spreads, turned into each team's implied pointsCached one hour
National NFL news coverageRecent reporting from the major football news outletsPer run; anything older than eight days is dropped

What we deliberately don't use: projection feeds, consensus rankings, ADP and tier lists. Those are the same numbers your leaguemates already have. Ranking and projection articles are filtered out of the news the agent reads, so conventional wisdom can't sneak in the back door.

A player is judged with his whole team, not alone

A running back is a bet on five linemen, a play-caller and a game script. Rankings flatten all of that into one number. The agent keeps it:

Where a fact doesn't exist, it stays blank rather than being guessed. Tracking data only covers players above weekly usage minimums, so a deep waiver flier simply has less to go on, and the call says so.

Where AI is used, and where it isn't

The split matters, because it's what keeps a confident-sounding model from inventing a statistic.

Plain code, no model

  • Reading and assembling every fact above
  • Working out which lineup slots are even in question
  • Guardrails, FAAB caps and confidence escalation
  • Grading calls after the games
  • Scheduling, delivery and billing

Claude

  • The start/sit judgement for slots genuinely in question
  • Ranking waiver targets and suggesting a bid
  • Evaluating a trade offer and drafting one worth proposing
  • Writing the Tuesday league report
  • Auditing recent calls against your philosophy for drift

Recommendations come back through structured tool calls with a fixed schema, so Claude returns data, not an essay, and every call is prompted to use only the supplied facts it was given.

Jev filters the noise before Claude ever sees it

Football coverage is mostly opinion: rankings, start/sit takes, waiver hype, and speculation about what a coach might do. Feeding that to a reasoning model is how you end up with recommendations that just echo the internet.

So every headline goes through Jev, a small, fast classifier, before anything reaches the reasoning step. Jev labels each item as real reporting or not: an actual injury, a confirmed role change or a transaction stays; analyst projections, rankings, ADP, tiers and speculation are dropped, as is anything older than eight days. Claude then reasons about what happened, not about what other people think will happen. It is also the cheap way round: classifying a headline costs a fraction of a reasoning call, so the expensive model only ever reads the survivors.

Lineup calls aren't a weekly re-shuffle of your whole roster. Only slots that are actually in question get reconsidered: an injury tag, or a slot your philosophy says to revisit by matchup, like flex and defense. Your studs stay in.

Three hard limits

These sit in front of every recommendation, regardless of what the reasoning concluded.

It grades itself, in public

Once a week is played, the agent goes back and scores its own calls against what happened.

Why this isn't another rankings site

Consensus tools answer "what does the crowd think about this player?" That's a fine question, and it's the same answer every manager in your league already has.

Consensus toolsFantasy Agent
The core productRankings averaged across expertsA decision for your roster, with its reasoning
Whose strategyThe average expert'sThe philosophy you wrote, per league
Player contextMostly the playerLine, role, efficiency, game environment, weather, news
When it disagreesAveraged awayFlagged as against conventional wisdom, with the case
AccountabilityRarely revisitedEvery lineup call graded against the alternative
Effort from youRead rankings, decide, repeatRead an email, make the move

The honest limit: this is an advisor, not an oracle. It won't tell you it's certain when it isn't, and it can't submit anything for you.

The data model, in detail

The questions worth asking before paying for anything like this.

What exactly is the underlying data model?

There's no single vendor model behind it. For each player and week, the agent builds a brief of public facts (usage, efficiency, line availability, weather, implied points, filtered news), combines that with your league's state and your written philosophy, and asks for a structured judgement. Facts are gathered deterministically in code; only the judgement is a model call. Briefs hold public facts only, with no roster or philosophy in them, which is why two managers with opposite philosophies can get opposite calls on the same player.

Which projection providers does it use?

None, deliberately. No projection feed is bought or scraped, and the agent doesn't publish weekly point projections. It models opportunity and situation instead: snap share, target and carry share, depth-chart role, tracking-data efficiency, offensive-line availability, implied team total and weather. Buying a projection feed would mean selling you a slightly rearranged version of what your leaguemates already see for free.

Which injury feeds does it use?

Three layers. Your league platform's injury designation, body part and notes for every rostered player, re-read on every run. Filtered headlines from the major national football outlets for the reporting behind a tag. And for the offensive line, whether each of a team's five most-used linemen is carrying a tag or missed the last game, which is the injury signal most tools skip entirely.

How frequently does data refresh?

Every run rebuilds anything stale, and nothing is served past its life: betting odds one hour, shared player briefs six hours, league data, weekly usage stats, the schedule, snap counts and tracking files up to a day, news dropped after eight days. A Tuesday forecast and a Tuesday closing total are both wrong by Sunday, so a Sunday lineup check refetches them rather than reusing Tuesday's.

How does it model player projections?

It doesn't produce a projected point total, and that's a design choice rather than a gap. A single week of fantasy points is mostly noise: a player can score twice on four touches and be a bad start next week. So the agent measures the volume a player is getting and its direction, adds the situation around him, and asks for a decision under your philosophy. You get "start him, and here's why" instead of a decimal that implies precision nobody has.

How does it handle uncertainty?

Four ways. Every call carries high, medium or low confidence. Low confidence is always escalated to you. Missing facts stay missing rather than being filled with a plausible guess, and the call is told what it doesn't know. And the confidence itself is audited: your dashboard charts hit rate by stated confidence, so if "high" isn't beating "low", it's visible rather than buried.

How does it calculate trade value?

There's no numeric trade-value chart. An offer is judged on roster fit, each player's current context and rest-of-season outlook, with an explicit instruction to apply your philosophy over generic trade-value consensus. The agent will also draft one offer worth proposing, with the case for why the other manager would say yes. Trades are advisory only and always routed to you; nothing is ever sent to your league.

How does it model waiver probability?

It doesn't predict whether your claim will win, because that depends on budgets and intentions no public API exposes. Candidates come from your league's trending-add counts, get ranked by how much they'd actually improve your roster (compared against your whole bench, not just the same position), and carry a suggested FAAB bid capped at 35% of your remaining budget, or zero in waiver-priority leagues.

Does it use expert consensus?

No. Consensus rankings, ADP and tiers aren't ingested at any point, and ranking or projection articles are filtered out of the news the agent reads. Instead, when a recommendation runs against conventional wisdom, it's flagged as such with a note explaining the disagreement. You see the contrarian call and its reasoning, rather than having it averaged into the middle.

Does it use betting-market information?

Yes, where betting-market data is configured. Game totals and spreads are converted into each team's implied points for the week, which is the cleanest available read on how much scoring a game is likely to hold. It's cached for an hour so every customer running in the same hour shares one fetch. Without a key the field is simply left empty rather than estimated.

How does it handle matchup data?

The NFL schedule and the forecast for uncovered stadiums come from public sources (your league platform publishes neither), and the game environment comes from implied team totals. Your own head-to-head opponent for the week comes from your league. Matchup is weighted where it genuinely swings a call, like defense and flex slots, and deliberately not used to bench a proven starter, because that's where matchup-chasing usually costs people weeks.

What models generate recommendations?

Claude Sonnet 5 from Anthropic makes every recommendation, through structured tool calls with a fixed schema, so answers come back as validated data rather than prose. Each prompt is restricted to the supplied facts and told not to invent injuries, news, scores, projections or rankings.

A second, much smaller model does the filtering. Jev is a fast classifier that reads every incoming headline and labels it real reporting or opinion: an injury, a confirmed role change or a transaction is kept, while analyst rankings, projections, ADP, tiers and speculation are thrown away. Only what survives reaches Claude, so recommendations rest on what happened rather than on what commentators predict, and because classifying a headline costs a fraction of reasoning about one, the filter also keeps the bill down. Independent calls within a run are then sent as one half-price batch.

How are recommendations backtested?

Automatically, against the only scoreboard that matters. Once a week is played, each lineup call is graded hit, miss or mixed by comparing the recommended starter's actual fantasy points against the player he'd have replaced, or against the best eligible bench option for a keep call. Waiver and trade calls record whether you followed them. Season hit rate and calibration by confidence are on your dashboard from the first graded week, and the sample is small early: a handful of graded calls is a hint, not a verdict.

See it on a real season

Watch a season replay call by call, or set your own agent up in a few minutes. $10/month, August through December only.