Blog

DOB Capital · September 21, 2026 · 12 min read

The Data Moat: Every Simulation Compounds Our Advantage (Build in Public #3)

How the DOB simulator builds a proprietary LATAM infrastructure risk dataset. 8 data points per simulation. The flywheel of predictive underwriting.

#build-in-public#data-moat#simulator#dataset#flywheel
Share
The Data Moat: Every Simulation Compounds Our Advantage (Build in Public #3)

The Data Moat: Every Simulation Compounds Our Advantage

Build in Public #3

This is the third post in our Build in Public series, where I share the honest, behind-the-scenes thinking that goes into building DOB Capital. Part 1 covered why we built a simulator before a product. Part 2 explained how we went from a spreadsheet to a multi-factor rate engine. This post is about the thing that keeps me up at night -- in a good way.

Every time someone runs a simulation on DOB Capital, they give us something that doesn't exist anywhere else in the world: a complete infrastructure operator risk profile for a Latin American market, with revealed economic preferences, submitted voluntarily.

Let me explain why that matters, why it compounds, and why it might be the most valuable thing we're building.

What We Capture (And Why Each Field Matters)

Every simulation captures 8 primary data points, plus several derived values. Here's exactly what they are and why each one is valuable.

The 8 Primary Inputs

1. Asset type (10 categories: data centers, energy, SaaS, real estate, fleet, mining, industrial, agriculture, health, other)

This tells us what kind of infrastructure is looking for financing in LATAM. Not what banks report in their portfolios. Not what government surveys say. What actual operators, sitting at their desks, are trying to finance right now.

2. Capital amount ($50K to $10M, 8 tiers)

Revealed demand. Not "how much would you like to borrow in theory" but "how much do you need for a specific project you're working on." The distribution of capital requests by country and asset type tells us where the real financing gaps are.

3. Term (1 to 10 years)

Combined with capital amount, this reveals the operator's view of their own project economics. A 2-year term on a $500K solar installation suggests different cash flow expectations than a 7-year term on the same amount. The term distribution by asset type gives us insight into how operators think about payback periods.

4. Country (12 options: Chile, Peru, Colombia, Mexico, Brazil, Ecuador, Costa Rica, Panama, Uruguay, Paraguay, Other LATAM, Outside LATAM)

Geographic demand distribution. Which markets have the most operators actively seeking alternative financing? Where is bank credit most broken? The answer isn't always what you'd expect from reading macro reports.

5. Collateral availability (binary: yes/no)

Does the operator have physical assets or guarantees to secure the financing? This single binary tells us about the operator's capital base and risk profile. Operators without collateral are the ones banks automatically reject -- and they're often running perfectly viable projects.

6. Credit history (binary: yes/no)

Has the operator been part of the formal financial system? In LATAM, where 70% of SMEs lack access to credit, a "no" here doesn't mean "bad borrower." It means "invisible to the banking system." This distinction is everything.

7. Operational track record (binary: day-1 distribution ready or not)

Can the asset generate revenue from day one? A solar installation with a signed PPA is day-1 ready. A greenfield project with no contracts is not. This tells us about the project's maturity and the operator's ability to de-risk.

8. Service contract (binary: yes/no)

Does the operator have a formal contract with an end customer? This is one of the strongest predictors of repayment capacity. An operator with a 5-year service contract has a fundamentally different risk profile than one selling into the spot market.

The Derived Values

From these 8 inputs, our engine calculates:

  • Risk score (0-7): A composite score that weights collateral, credit history, track record, distribution readiness, and service contracts.
  • Bank access difficulty (easy/medium/hard): Based on country-specific banking criteria and the operator's profile. This tells us whether the bank would even consider the application.
  • Indicative rate: The operator's estimated financing cost through our platform.
  • Potential rate: The rate achievable after due diligence verification.
  • Bank comparison rate: What the operator would pay at a bank -- if the bank said yes.
  • Rate differential: The spread between DOB and bank rates, which tells us how competitive our offering is for each specific profile.
  • Monthly payment: Interest-only monthly payment estimate.
  • Funding preference (crypto/fiat): Revealed preference for settlement infrastructure -- captured at lead submission.

That's 8 primary inputs plus 8 derived values. Sixteen data points per simulation. Each one structured, comparable, and compoundable.

Why This Dataset Doesn't Exist

Let me be very specific about why this matters. This dataset -- individual-level infrastructure operator risk profiles across LATAM -- does not exist anywhere in the world. Here's who might have pieces of it, and why they don't have the whole picture:

Banks have credit data on the operators they approve. But banks reject 70% of PYME applications in LATAM. Their dataset is survival-biased -- it only contains information about the operators who already had access. The invisible 70% is exactly where the opportunity lives, and banks have no data on them.

Central banks publish aggregate statistics. The BCRP in Peru publishes average PYME rates. The BCB in Brazil publishes credit volume by sector. The Banco de Mexico publishes intermediation margins. These aggregates are useful for macro analysis but tell you nothing about individual operator profiles, specific asset types, or revealed capital demand at the project level.

Credit bureaus (Equifax LATAM, TransUnion, DICOM in Chile) have payment history data, but only for operators who already participate in the formal credit system. Same survival bias as banks. And their data doesn't include asset-level information -- they know if someone paid their credit card bill, not whether their solar installation generates $40K/month.

Development banks (IDB, CAF, BNDES, CORFO) have project-level data for the deals they finance. But development bank financing serves a tiny fraction of the market. BNDES, the largest development bank in Latin America, reaches approximately 2% of Brazilian SMEs. Their dataset is deep but narrow.

Fintech lenders in LATAM (Credijusto, Konfio, Cora, etc.) have lending data, but focused on working capital and invoice factoring -- short-term, high-frequency lending. They don't capture infrastructure project profiles with multi-year terms and asset-specific risk factors.

Government surveys (Sebrae in Brazil, CORFO in Chile) periodically survey SME financing needs. These are valuable but infrequent, self-reported, and don't capture the real-time, revealed-preference data that comes from someone actively trying to finance a specific project.

DOB's simulator sits in the gap between all of these. It captures data from operators who are actively seeking financing (revealed preference, not stated), across the full spectrum of bank access (including the 70% who get rejected), at the individual project level (not aggregated), in real time (not periodic surveys), and structured in a format that enables cross-country, cross-industry comparison.

The Flywheel

Here's where it gets interesting. This dataset isn't static. It compounds.

More simulations lead to better rate calibration. As we accumulate data, we can validate and refine our rate model. Are operators in Peru's energy sector consistently requesting 5-year terms? That tells us something about PPA structures in Peru that should influence our term-risk adjustment. Are Chilean SaaS operators disproportionately lacking collateral? That might mean our collateral weight is too punitive for software businesses in Chile.

Better rate calibration leads to more accurate results. When an operator runs a simulation and gets a rate that aligns with their market reality, they trust the tool. When the bank comparison is accurate -- when the operator thinks "yes, that's exactly what my bank quoted me" or "yes, that's exactly why my bank said no" -- credibility compounds.

More accurate results lead to more trust. Operators share tools that give them useful information. "Hey, have you tried this simulator? It told me my rate would be X and it was spot-on" is the most powerful growth mechanism in B2B fintech. It costs nothing, can't be faked, and compounds with every accurate simulation.

More trust leads to more simulations. And the cycle repeats.

This is a classic data flywheel. But unlike consumer data flywheels (which have been talked about to death), this one operates in a market where the baseline is near-zero data availability. In consumer fintech, you're competing against companies with billions of transaction data points. In LATAM infrastructure credit, you're competing against a vacuum.

What We Can See (And What We'll Eventually See)

Today, with our current dataset, we can answer questions like:

  • What's the average capital requested by energy operators in Peru?
  • What percentage of Colombian operators have collateral available?
  • Which countries have the highest concentration of operators who prefer crypto settlement?
  • What's the most common risk score by asset type?

These are useful for calibrating our rate model and understanding our market. But they're descriptive. They tell us what the market looks like today.

With scale, the dataset becomes predictive:

  • Conversion prediction: Which operator profiles are most likely to proceed from simulation to account creation to asset publication? If we can identify the characteristics of operators who convert, we can optimize the entire funnel -- not with marketing tricks, but with better product design.
  • Market heat mapping: Where is demand growing fastest? If simulations from Uruguayan agriculture operators tripled last quarter, that's a leading indicator of market opportunity that no macro report would capture for another 6-12 months.
  • Default prediction: Eventually, when we have financing history, we can correlate pre-financing profiles with repayment outcomes. Which combination of risk factors actually predicts default? The answer probably isn't what traditional credit models assume.
  • Rate optimization: Dynamic rate adjustment based on accumulated performance data, not just macro indicators. The rate for a Chilean SaaS company with collateral and a service contract should be informed by how similar profiles have performed, not just by Chile's policy rate and a theoretical spread model.

This is the long game. We're not there yet. The dataset is young. But every simulation moves us closer.

The Privacy Architecture

I want to be direct about how we handle this data, because trust is the foundation of the flywheel.

For analytics: All data is anonymized. Country-level and industry-level aggregates are used for rate model calibration and market analysis. No individual profiles are used in aggregate reporting.

For the operator: Your simulation data is yours. It's accessible only to you within the platform, and (with your explicit consent) to DOB for the purpose of structuring your financing. We don't sell individual profiles. We don't share them with third parties. We don't use your specific simulation to market to your competitors.

For the dataset: The compound value comes from patterns, not individuals. We need to know that "energy operators in Peru typically request $500K-$1M over 5-year terms with a risk score of 4-5." We don't need to know that you specifically requested $750K.

This isn't just ethical -- it's strategic. An operator who doesn't trust the privacy of the simulator won't use it. An operator who doesn't use it doesn't contribute to the dataset. Breaking trust breaks the flywheel.

An Honest Assessment

Let me be honest about where we are. The dataset is small. We're early. The flywheel has started turning but it hasn't reached escape velocity.

We don't have enough data yet for meaningful predictive models. Our rate calibration is informed more by macro research (OECD data, central bank publications, EMBI spreads) than by accumulated simulation data. The market heat mapping is interesting but not yet statistically significant.

What we do have:

  • Structure. Every data point is captured in a consistent, structured format from day one. When the dataset is large enough for machine learning, it won't need to be cleaned or normalized -- it was designed for analysis from the start.
  • Real-time capture. Unlike surveys or government reports, our data reflects what operators are doing right now, not what they did 18 months ago.
  • Global comparability. The same 8 inputs, the same scoring model, across 12 countries and 10 asset types. That's 840 possible combinations, each one a micro-market that can be analyzed independently or compared across dimensions.
  • Compounding. Every simulation makes the dataset more valuable. There is no point at which additional data stops being useful. The marginal value of each simulation increases as the dataset grows, because each new data point can be compared against more existing patterns.

The simulator isn't just a lead generation tool. It isn't even primarily a lead generation tool. It's the foundation of a credit intelligence platform that, at scale, will know more about LATAM infrastructure credit demand than any bank, central bank, or development institution.

That's the moat. And every simulation digs it deeper.

Your Simulation Is Part of the Picture

If you're an infrastructure operator in Latin America -- whether you're running solar panels in Chile, managing a vehicle fleet in Mexico, or operating a data center in Brazil -- your simulation contributes to something larger than your individual rate estimate.

You get immediate value: a rate, a bank comparison, a clear picture of your financing options. But you also become part of a dataset that, over time, will make alternative financing more accurate, more affordable, and more accessible for every operator in the region.

The first step takes 2 minutes.

Simulate your rate now -- free, confidential, no commitment. See where you stand, and help build the dataset that changes how LATAM infrastructure gets financed.

This is Build in Public #3. Previous: From Spreadsheet to Rate Engine (Build in Public #2). This is post 19 of our 20-part series. Next: The $250 Billion Gap: Closing It Takes Rails, Not Rhetoric.

DOB Capital

DOB Capital

Alternative financing for asset operators in LATAM. No banks, no borders.

Simulate