TLDR: noticed consistently beat flagship models at 1/8 the cost and achieved 2.8x the performance of comparably priced models. View the results.
This post is a work in progress. We’re continuing to run benchmarks and will update the results as testing progresses.
When a founder puts their trust in an AI assistant to help plan a fundraising week, it must answer a deceptively difficult query: “which relationships should we act on now?”.
A CRM stores a customer’s identity, related tasks and their deadlines. Meeting transcripts record a point-in-time state of a relationship and a set of next steps. A calendar event can show that the follow-up has already been arranged and a one-line email can completely change everything up to the point when it was received. Correctly answering the founder’s query depends on reconciling these records to first define, if nothing else, what is a relevant person in the current context.
This technical write-up describes how we built noticed’s first benchmark for evaluating an AI agent’s ability to plan next steps in founder-led fundraising and sales. We tested whether noticed could turn fragmented relationship evidence into better next-step recommendations than a general-purpose AI agent using a set of external MCP integrations. In this benchmark, noticed consistently beat flagship models at 1/8 the cost and achieved 2.8x the performance of comparably priced models.
Generating synthetic data
We’re currently working with our first research partners. Building this benchmark is the first step in a plan to design and build noticed’s proprietary AI harness and models to cover a set of engineering challenges such as data enrichment, identity matching, aggregation and generalized inference. These models will be trained on our proprietary data and customer interactions recorded by our system.
For the first version of our benchmark, we chose to replicate two business contexts and generate synthetic data based on life-like scenarios. Kilnbeam has two founders building quoting and capacity-planning software for precision manufacturers. It is raising €900,000 with three unpaid design partners. Docklane has two founders, two operators and five paying customers for its freight-exception software. It needs to refine its ideal customer profile (ICP) and find five more customers.
Each team received five tasks, covering selection, introductions, meeting preparation, follow-through and coordination. These tasks require decisions across relationships: who fits, who can introduce them, what is already known, what needs to be done, and who should act. The two worlds cover founder-led fundraising and customer acquisition; they do not test hiring, broader networking or relationship outcomes. These are development cases; any similarity to actual persons, living or dead, or actual events, is purely coincidental.
We authored both sides of each world, including investor mandates, customer adoption problems, buying constraints and introduction conditions. The shared provider fixtures contain 65 people records, 63 company records, 34 meetings, 19 CRM notes, 25 deals and 12 tasks. The additional fixtures contain 70 emails, 43 WhatsApp messages and 35 calendar events, plus 65 LinkedIn profiles and 65 connection records for the same people. The fixtures were generated by a deterministic Python authoring script, with fixed fictional records and stable UUIDv5 identifiers; it makes no model or network calls and uses no customer data. The freeze checks verify rubric evidence references, meeting IDs and owners, task coverage and 100-point totals. Dates and conflicting records are authored explicitly; these checks do not establish a complete chronology audit or independent human review.
Aside from the synthetic datasets, added contextual challenges were set up to mimic real-world situations where interest is not a final decision, a customer result is not necessarily approved sales material and a newer message may change a deadline while leaving the original assignee and obligation intact. Time has also been added as part of this simulation and the evaluation clock was fixed at September 8, 2026, 09:00 UTC.
Here’s a request sample and some of the synthetic data that was generated for this scenario:
“What fundraising follow-ups should we do in the next seven days, and which should we postpone? Give owners and draft the two most useful messages.”
| Record | What it establishes |
|---|---|
| Granola meeting, August 28 | Sofia wants evidence that Kilnbeam's workflows are repeatable across manufacturers. Leon will use synthetic examples. |
| Attio task, created September 1 | Leon should follow up with Sofia by September 9. The task is incomplete. |
| Gmail message, September 6 | Sofia is away until September 15, asks not to be chased, and welcomes an update on September 16. |
One investor, Sofia Malik at Cedarline, has an open follow-up task. These three records determine what should happen. The expected action is to prepare the update and defer contact until September 16. The email changes the contact date; it does not cancel the work or change its owner.
Besides the final responses, evaluation records retain the tool calls and records returned to the model: speakers, task assignees, linked records, deadlines, completion flags and attendee responses, all referenced by source ID during the grading stage of the benchmark.
Benchmark design
We run the same tasks through two harnesses: the baseline harness connects a general-purpose model to locally simulated Granola and Attio MCP tools; the noticed harness runs noticed locally, using its own instructions and native tools over imported relationship data. The baseline was used to test model candidates while retaining the task prompts, tool definitions, source fixtures and judging criteria. noticed imports the same Granola and Attio records, plus Gmail, WhatsApp, Google Calendar and LinkedIn data, and retrieves them through its own relationship history and profile tools.
This setup allows us to compare the resulting workflows: the baseline can read full transcripts and CRM notes whereas noticed receives imported summaries, metadata and derived facts, including evidence unavailable through the baseline’s two integrations. Both harnesses receive the same user task and neither answering model receives the private judging criteria. The simulations expose meeting lists, summaries and transcripts, CRM records and notes, tasks and workspace members.
| Configuration | Mini baseline | Sol baseline | noticed |
|---|---|---|---|
| Answer model | gpt-5-mini | openai/gpt-5.6-sol | undisclosed |
| Reasoning | low | low | medium |
| Shared Granola/Attio fixtures | yes | yes | yes |
| Full transcripts at answer time | yes | yes | no; imported summaries and derived facts |
| Additional sources | none | none | Gmail, WhatsApp, Calendar, LinkedIn |
| Maximum tool steps | 12 | 12 | 12 |
| Setup | local provider fixtures | local provider fixtures | local import, identity reconciliation and fact extraction |
noticed’s source-fact extraction setup is separate from answer generation and excluded from its cost. These differences mean the comparison measures the combined effects of source access, representation, model and harness.
The ten cases below cover Kilnbeam’s fundraising work (K1–K5) and Docklane’s customer acquisition work (D1–D5). The criteria summarize the original source-backed rubrics; each case is scored out of 100, with critical violations tracked separately.
| Test case | Task | Judging criteria summary |
|---|---|---|
| K1 · Investor shortlist | Choose five investors to prioritize this week, with introduction paths and teammate owners. | • Stage, geography, cheque size and pre-revenue fit • Exclude unsuitable, paused or declined funds • Supported routes, correct identities and owners • Current timing; distinguish interest from commitment |
| K2 · Introduction request | Prepare the strongest introduction request and the work needed before it can be forwarded. | • Leon → Ana → Edda route and relationship evidence • Deliver promised demo before forwarding • Double opt-in and accurate round/pilot claims • Protect private information; distinguish fixed bugs from untested capacity |
| K3 · Investor meeting preparation | Brief both founders for tomorrow’s meeting, including roles, concerns and questions. | • Confirm correct meeting, attendees and identities • Relevant history and commercial/technical speaking roles • Honest traction and technical limitations • Safe demo materials and useful diligence questions |
| K4 · Fundraising follow-through | Plan the next seven days of follow-ups, defer inappropriate contact and draft two messages. | • Prioritize outstanding promises with owners and dates • Respect leave, review windows, pauses and passes • Avoid duplicate asks or reopening completed work • Two useful drafts; no claimed outreach or invented outcomes |
| K5 · Founder pipeline coordination | Reconcile the fundraising pipeline into one coordinated plan for both founders. | • Separate active, waiting, deferred and declined prospects • One owner per investor; preserve introducer provenance • Resolve same-name people and stale CRM state • Dated next steps and prerequisites; no invented committed capital |
| D1 · ICP refinement | Use the first five customers to propose a narrower ICP, exclusions and validation questions. | • Workflow, volume, buyer and integration fit • Correct customer metrics, denominators and attribution • Account for adoption, support burden and weak-fit customers • Treat five customers and before/after results as provisional evidence |
| D2 · Next five prospects | Select five new accounts with buyers, qualification gaps, introduction paths and owners. | • Five distinct suitable prospects, excluding current customers • Evidence of fit; explicit budget/integration gaps • Distinguish champions from decision-makers • Correct identities, owners and consent; respect opt-outs and pauses |
| D3 · Re-engagement timing | Choose prospects ready to re-engage this week and draft the two best messages. | • Recent, evidenced triggers for contact • Correct owner, recipient and permitted channel • Respect migration delays, pauses and opt-outs • Specific drafts with bounded asks; no invented history or confirmed outcomes |
| D4 · Customer referrals | Choose a customer advocate and referral target; prepare an internal plan and customer-facing request. | • Pine → Elm as the permissioned immediate route • Separate reference permission from introduction consent • Defer asks while support issues or timing barriers remain • Ground advocacy in adoption; protect private data and avoid invented endorsements |
| D5 · Coordinated acquisition plan | Give all four teammates a two-week plan to pursue five more customers. | • Assign roles while retaining relationship owners • Sequence work around triggers, dependencies and dates • Explicit stop conditions for fit, consent, budget and support • Avoid duplicate outreach; protect customer-success capacity; no guaranteed sales |
Scoring and judging
Each stack answered ten scenarios three times, producing 30 answers per stack (90 in total). The judging model received each answer, its retrieval trace, team description, source-access description, complete source corpus and private rubric. It returned an award, rationale, evidence IDs and failure classification for every criterion. Code checks that each expected criterion appears exactly once and receives zero, half or full credit. The judge distinguishes inaccessible sources, retrieval failures, reasoning failures and incomplete answers. Inaccessible evidence can still cost points under the unchanged 100-point rubric; there is no adjustment for each system’s attainable score. We have not reported a per-system failure-type breakdown here, so the overall scores combine source coverage with retrieval and reasoning quality. Missing evidence alone is not treated as fabrication.
For example, 20 of the fundraising query's 100 points cover timing and completed work: “defer Sofia until September 16”, “respect another investor's review date”, and “recognize an already-sent deck”. The remaining criteria cover outstanding commitments, owners, pauses, two useful drafts, unresolved work and dated evidence.
Critical violations remain separate from the numeric total. A high score cannot make contact during an explicit no-contact window acceptable. Missing judge outputs and setup failures stay incomplete rather than receiving invented zero scores.
scenario_score = median(repeat_1, repeat_2, repeat_3)
overall_score = median(the 10 scenario_scores)For noticed, the middle two scenario medians are 90 and 92.5, giving 91.25. This is a rubric score, not a percentage of successful tasks or a measured business outcome.
Results
Benchmark score vs. cost per answer
Overall rubric score versus average generation cost per answer (USD). Logarithmic cost axis.
Swipe the chart horizontally to see all labels.
noticed achieved 1.52× Sol’s score at 1/8.25 of its generation cost. Mini remained the cheapest stack.
| System | Score / 100 | Three repeat scores |
|---|---|---|
| noticed | 91.25 | All 10 scenarios |
| GPT-5 Mini | 32.5 | All 10 scenarios |
| GPT-5.6 Sol | 60 | All 10 scenarios |
| System | Overall rubric score / 100 | Median answer time | Generation cost, 30 answers (USD) | Listed input / output rate (USD per 1M tokens) |
|---|---|---|---|---|
| GPT-5 Mini + Granola + Attio | 32.5 | 29.07s | $0.215 | $0.25 / $2.00 |
| GPT-5.6 Sol + Granola + Attio | 60.0 | 39.13s | $5.095 | $2.00 / $10.00 |
| noticed | 91.25 | 104.35s | $0.618 | $0.07 / $0.15 |
Table 1. Testing is continuing with other models; this table contains the three evaluated systems and will be updated as testing continues.
Generation cost includes attributed noticed helper calls, but excludes ingestion and extraction setup, judging, development iterations, hosting and subscriptions. It is the marginal cost of answering this suite, not the total cost of operating the product.
Mini’s $0.215 is a token-based estimate using OpenAI rates recorded on September 8, 2026, with cached input charged at the full input rate. Sol’s $5.095 and noticed’s $0.618 are USD costs returned by Vercel AI Gateway for the September 14 and September 10 runs, respectively; they are not the ledger’s padded budget-reservation units. Helper calls are attributed to noticed only within each answer’s exclusive execution window.
Across the 30 answers, Mini used 271,723 input and 73,320 output tokens; Sol used 855,788 and 79,845; noticed’s answer model used 3,198,844 and 295,218. noticed also used 8,689 input and 948 output tokens for language-model helpers, plus 55 embedding tokens. Output usage includes reasoning tokens. Gateway costs retain provider-reported charging; we have not isolated cache savings.
Answer time runs from local answer orchestration to the final response, including model execution, tools, helpers and any recovery. Runs did not overlap within each system’s suite; cache warmth was not controlled. No saved answer used the extra reply-recovery pass. The slowest answers took 73.67 seconds for Mini, 58.33 for Sol and 224.94 for noticed. Setup and judging are excluded, and no amortized setup-cost analysis is included. Provider retries are metered per call but are not broken out separately here.
Professional relationship benchmark
Median of ten scenario medians. Each scenario has three answers, graded on a 100-point rubric. Higher is better.
Swipe the chart horizontally to see all labels.
noticed: 91.25 · GPT-5.6 Sol: 60 · GPT-5 Mini: 32.5
Synthetic workflow comparison, not a model-only leaderboard. noticed has additional source access. A score is not a task success rate.
| System | Score / 100 | Three repeat scores |
|---|---|---|
| noticed | 91.25 | All 10 scenarios |
| GPT-5 Mini | 32.5 | All 10 scenarios |
| GPT-5.6 Sol | 60 | All 10 scenarios |
noticed achieved 1.52× Sol's score at 1/8.25 of its generation cost. It achieved 2.81× Mini's score at 2.88× its cost.
Against Sol, the largest median gaps were investor meeting preparation (+65 points) and ICP refinement (+57.5); the next-five-prospects task was much closer (+5). noticed led in nine of ten scenarios. Sol led on the introduction task. That task (K2: noticed 30, Sol 60) has a known fixture–rubric conflict: the expected introduction between two synthetic people coexists with an already-booked meeting with one of them. We retain the historical scores; resolving the conflict requires a new benchmark version and a rerun of all systems.
Score variation across repeated runs
Average highest-to-lowest score gap within each scenario, across three repeated answers. Lower is better.
Swipe the chart horizontally to see all labels.
noticed had a smaller score range in 8/10 scenarios versus Mini (1 tie), and 7/10 versus Sol.
Three repeats are descriptive, not a reliability guarantee. Judge variation contributes to the spread; scores near 100 have less room to vary upward.
| System | Score gap | Three repeat scores |
|---|---|---|
| noticed | 10 | All 10 scenarios |
| GPT-5 Mini | 22.5 | All 10 scenarios |
| GPT-5.6 Sol | 26 | All 10 scenarios |
Repeated answers were also more consistent in score. For each scenario, we measured the gap between the highest and lowest of its three scores, then averaged those gaps across the ten scenarios: 10 points for noticed, 22.5 for GPT-5 Mini and 26 for GPT-5.6 Sol. noticed had a smaller spread in eight scenarios versus Mini (one tie) and seven versus Sol. Judge variation may contribute to the spread, and scores near the 100-point ceiling have less room to vary upward. As for latency, noticed took about 104 seconds per answer, compared with 39 for Sol. These descriptive results from ten tasks across two synthetic worlds do not yet establish statistical significance; the three repeats per task are not independent samples of relationship decisions.
Next steps
Early results in this benchmark suggest the noticed harness performs better at the given tasks than other, more expensive models in a general-purpose harness. Nevertheless, the tasks are limited in scope and the testing is limited in volume and model candidates. This technical write-up is a WIP and will continue to be updated as more testing is conducted. It remains relevant to find better answers for these questions:
- Same model, different harnesses: how much comes from the system?
- Same source access, different retrieval and representation: how much comes from assembling context?
- Unseen worlds and an independent judge: does the advantage generalize?
The research team at noticed is also interested in individually benchmarking the several components that make up our system, namely our data enrichment and identity matching solutions, against other alternatives on the market. Testing each of these individually will provide a more complete picture of whether, where and why our system outperforms alternatives. It will also highlight the path for improvement of generalized benchmarks.
It’s our belief that by working with more research partners and creating better synthetic data from the experiences we learn from them, we’ll be able to create better, more demanding benchmarks that test whether noticed can excel at building professional relationships.
![[WIP] Building a benchmark for professional relationship decisions](/_next/image?url=%2Fblog%2Fimages%2Frelationship-benchmark-atmosphere.png&w=3840&q=75&dpl=dpl_8r7AKgngPkQK5gMB8uv3CP9cUigw)