Skip to main content
PrismCV
JobsExtensionPricing
LoginCheck Your Resume
Check Your Resume
← Back to all jobs

Founder Fellow - Applied AI

Pennant · New York, NY, US

On-site
Mid
Check your resume against this jobApply on Ycwaas

Job Description

Internship, seasonal with potential to convert to full-time. New York City, in-person required. Reports to the Co-founder & CTO and works directly with both founders and domain experts.

About Pennant

Pennant is a YC-backed company building software for corporate governance, starting with proxy voting and company engagement.

Institutional investors, public companies and their advisors make consequential decisions using information scattered across filings, policies, research and conversations. We bring that information together so teams can understand the evidence, apply their own judgment and preserve why they made a decision.

Our ambition is a world model for corporate governance: a system that connects institutional knowledge, policies, decisions and outcomes. Getting there starts with reliable data and software customers trust in their daily work.

The fellowship

Make our document understanding measurably better and prove it.

Pennant turns dense public filings and customer documents into structured information that analysts rely on. Every number we show traces back to that step, so the interesting work is not writing the first prompt. It is finding where the system is wrong, deciding whether the fix belongs in code, in the prompt or in the evaluation set, and shipping the change with a regression test that would catch it next time.

You will join a team of two founders and our first engineering hires. Your work goes to customers who use it to make real governance decisions. We will hand you bounded problems with a clear success measure, teach you the filings, and expect you to carry each one through to a merged, tested change.

What you'll work on

  • Extend our evaluation harness. Add human-reviewed reference cases, measure completeness and consequential errors rather than aggregate agreement, and turn production failures into regression tests.
  • Diagnose extraction failures on real filings: a value read from the wrong table, a figure pulled from the wrong period, an item matched to the wrong record. Trace each to its cause and fix it where it belongs.
  • Build deterministic checks around model output so malformed or contradictory results are caught before anyone relies on them.
  • Run cost and quality experiments across models, prompts and caching, with numbers a founder can make a release decision on.
  • Write small tools for domain reviewers so labeling and spot-checks take minutes instead of hours.

Problems you might tackle

  • A model update raises average accuracy on our reference set but silently drops one field on a class of filings. Find the class, quantify it, and decide whether the update ships.
  • A filing uses a table layout we have never seen. Decide whether to extend the parser, adjust the prompt or add a targeted example, and show the change does not regress the rest of the corpus.
  • Our reviewers disagree with each other on a category of cases. Design the reference set so the ambiguity is explicit instead of averaged away.

What you bring

  • Solid programming in TypeScript or Python and comfort with SQL. Most of our code is TypeScript; Python is welcome for analysis.
  • Hands-on experience building something with LLM APIs that other people used, and a clear account of where it failed and what you did about it.
  • Care about measurement. You reach for a test set and a baseline before you reach for a bigger model, and you distrust a metric you cannot explain.
  • Patience with dense documents. You are willing to read a proxy statement closely enough to know when the model is wrong.
  • The habit of finishing: scoped change, tests, review, merge, and a short note on what you learned.
  • Availability to work in person in New York for the term.

Governance and finance knowledge are not prerequisites; we will teach the domain. Experience with information extraction, document parsing, financial or legal text, or LLM observability tools is useful. So is anything you have shipped on your own: repos, a product with users, a research project with results you can defend.

Our stack

A TypeScript monorepo with NestJS services, Go for some backend services, Postgres and BigQuery on Google Cloud. We use Anthropic and OpenAI models with LangSmith tracing, CI enforces coverage gates, and we use AI coding tools daily.

How we work

We work in person in New York and stay close to customers. We prototype quickly, use AI development tools where they help, and remain responsible for what we ship.

We narrow scope before compromising correctness, permissions or customer trust. We test representative cases and failure paths, observe what happens after release, and flag risks early. Fellows get real ownership of bounded problems and are expected to make their reasoning understandable to the team.

Apply

Send a short note and a link to something you built with a language model. Tell us one specific way it was wrong, how you found out, and what you changed. We would rather read that than a list of frameworks.

See how well your resume matches this job before you apply

Run a free ATS check