Proximal generates coding tasks with AI, surpasses $200M in revenue

1 day ago 8

Proximal, a San Francisco-based AI research lab, has reportedly crossed $200 million in annual revenue by using AI models to generate coding tasks at scale. The company’s approach leans on synthetic data generation and reinforcement learning to build what it describes as high-fidelity environments for training autonomous coding agents.

For a company founded in 2025 with roughly 25 employees, that revenue figure would be remarkable. It would place Proximal in rare company among AI startups that have scaled revenue faster than most enterprise software firms in history.

What Proximal actually does

Proximal builds automated pipelines that generate synthetic code, create evaluation rubrics, run quality checks, and simulate extended software development tasks using multi-agent orchestration.

The company released FrontierSWE, a benchmark containing ultra-long-horizon tasks meant to evaluate how well coding agents handle complex, multi-step engineering work. Even leading models like Anthropic’s Claude performed poorly on these extended benchmarks, suggesting that current AI coding tools still have significant ground to cover before they can autonomously handle serious software engineering.

Proximal operates out of both San Francisco and Bangalore, maintaining a lean team that reflects its engineering-driven philosophy. The company is backed by Scribble Ventures, with angel investments from individuals connected to OpenAI, Anthropic, and other prominent AI firms.

The revenue question

The $200 million revenue claim deserves some context. Proximal is currently classified as an early-stage, seed-backed enterprise. No public filings, company profiles, or industry directories have independently verified that figure. The company has not disclosed annual recurring revenue numbers through any publicly available channels.

For comparison, Proximal operates in the same broad category as companies like Cursor and Cognition, both of which have reported strong financial performances in the AI coding space.

Why synthetic data matters right now

The broader AI industry is running into a wall. Training frontier models requires enormous amounts of high-quality data, and the supply of human-generated code suitable for training is finite. Synthetic data generation, the process of using AI to create training data for other AI systems, has emerged as one of the most promising solutions to this bottleneck.

The FrontierSWE benchmark reflects this ambition. By creating tasks that require sustained reasoning over long time horizons, Proximal is effectively building the obstacle course that next-generation coding agents will need to master. The fact that current state-of-the-art models struggle on these benchmarks tells us something important about where the industry stands: the gap between “impressive demo” and “reliable autonomous engineer” remains wide.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article