Skip to main content

Propexo Connect is now available. 125 managed connectors from your operational stack into the warehouse you own. Read the announcement

AI readiness

Your Pipelines Are Landing Data. Your AI Still Can't Answer the Question.

Multifamily operators have solved data movement and not the layer above it. What AI needs is not more pipelines but a shared definition of what the data means.

Remen Okoruwa, Co-Founder & CEO, Propexo

Remen Okoruwa · Co-Founder & CEO, Propexo

8 min read

A staggered grid of connector tiles from the Propexo Connect catalog: property management systems such as Yardi, Entrata, RealPage, Rent Manager, Buildium, DoorLoop and MRI, operational tools such as HappyCo, EliseAI, Funnel, Matterport, ButterflyMX, Leonardo247 and Brivo, and warehouse destinations including Snowflake, Google BigQuery, Databricks and PostgreSQL.

One customer came to us already paying for a horizontal data platform, with the warehouse, the budget, and the tooling in place. Their own engineers were still hand-maintaining two pipelines, because that platform carried no native connector for the property management system at the center of their stack, or for the AI leasing tool that had become the busiest surface in their funnel. The two systems that mattered most were the two nobody else would carry.

That is the state of multifamily data in one scene, and it is why the industry’s data conversation is stuck. Moving data is close to a solved problem. Knowing what the data means once it has moved is not, and it is the part that decides whether any of the AI now being sold into this industry works.

Why moving the data stopped being the hard part

The last three years went into plumbing, and the plumbing largely got built. Yardi shipped an MCP connector while Entrata overhauled its API, AppFolio built out a marketplace, and most of the vendors that spent a decade resisting have stopped resisting. Every serious proptech system now has some way to get data out.

But despite all of it, the questions operators put to their own data are no easier to answer than they were. Ask how many units turned last quarter across a portfolio that runs three property management systems, and you get three answers, because each system counts a turn from a different event. Or take a harder one, which vendors touched a unit before a resident churned, and nobody can assemble the answer without a week of manual reconciliation.

The reason is not missing pipelines. It is that the systems on either end of those pipelines were never designed to agree with each other, and moving their data into one warehouse does not make them agree. All it does is put the disagreement in one place, which is progress, and is not the same thing as an answer.

What a year of connectors leaves behind

Open a multifamily warehouse that has been fed by connectors for a year and what you find is more awkward than a data gap. The same resident sits in it under four spellings, because the screening vendor, the payments processor, the access system, and the PMS each captured the name at a different moment. A unit’s identifier changes shape somewhere between the maintenance tool and the accounting ledger, and “occupancy” turns out to be defined four ways, none of them wrong inside its own system and no two of them the same.

Every count built across those systems is a guess wearing the costume of a fact. Pick a number your team reports monthly that crosses two vendors, and ask which system’s definition of the shared field won. If nobody can answer in under a minute, the number is a guess.

This is where the AI conversation and the data conversation stop being separate. An agent asked to flag at-risk renewals has to know which records describe the same resident, how a lease relates to a unit and a unit to a property, and what “at risk” means in terms the data can express. Skip any of the three and it still answers, fluently and confidently, and the answer is wrong in a way nobody catches for a quarter.

The narrow-workflow trap

Most of the market has responded by going narrow. An AI leasing assistant lives inside one system, so it can be trained on one schema and it works. A maintenance triage tool reads one work-order table and does a useful job. Each of these is real value.

The trap is that narrow is where they stop. A leasing assistant cannot tell you that the prospect it just qualified is the same person your screening vendor declined two months ago, because it cannot see the screening vendor. Nor can a triage tool weigh a work order against the renewal risk of the resident who filed it. The moment a question crosses two systems, the tool that was trained on one goes quiet, and crossing systems is where most of the operating leverage in this business lives.

Vendors went narrow because the alternative did not exist. Building an agent that reasons across the operational stack requires a shared definition of the stack, and nobody had one. Going narrow was the only way to ship something that worked at all. What has changed is that the ceiling is now visible, and the operators pushing hardest on AI are the ones hitting it first.

What you get when the layer above the pipe exists

A number you can take to your owner without caveating it. Right now a portfolio-wide occupancy or turn figure carries an invisible asterisk, because each system counted it from a different event and someone chose which one won. When every system agrees what the word means, the asterisk goes away, and the number survives the question “where did this come from” in a board meeting.

An answer to a question that crosses two departments. Which vendors touched a unit before that resident churned, or whether the properties with the slowest maintenance response are the ones losing renewals. Those questions die today not because the data is missing but because nothing writes down how a property, a unit, a lease, a resident and a work order relate to each other. That knowledge lives in one analyst’s head, which is why it leaves when they do.

And counts that are not quietly inflated. The same resident shows up in five systems under five spellings, so every population number you report is wrong by an amount nobody can quantify. Getting that right is unglamorous and it is the difference between a renewal-risk model you would act on and one you would not.

The industry has names for the machinery behind all of that: defined semantics, a real estate ontology, and entity resolution. Together they are an intelligence layer, and an AI agent is only ever as trustworthy as its grasp of what a word means, how things relate, and which records describe the same thing.

What we are announcing, and what is true today

What exists today is the pipe, and as of today it is available. Propexo Connect moves data from the multifamily operational stack into a warehouse the operator owns and controls, across 125 managed connectors spanning 23 categories of that stack. Every connector delivers raw extracts. For the ten property management systems our Unified API covers, you can take a normalized layer instead, or both. The rest of the stack lands raw, which is the boundary to know before you plan a model on top of it. We own authentication, pagination, rate limits, and schema drift, so a vendor’s API change is our problem rather than a sprint your team did not plan for. For the owner who wants that in a number, our own comparison pages put a custom connector at four to twelve engineer-weeks, and eight to fifteen for one you would trust in production. Multiply that by the number of systems your portfolio runs and you have the line item this replaces. Onboarding runs two to six weeks, varying with how ready your own data is and how many source systems are in scope.

The Propexo Connect stream selection screen for one connector: 23 accessible streams such as accounts, properties, inspections and form answers, each with its sync mode (incremental or full refresh) and cursor field.
Selecting streams on one Connect source: each stream shows its sync mode and cursor field.

The intelligence layer is the part still ahead. What exists now is its foundation, and today that foundation gets a name: we are announcing Connect Intelligence, an owner-controlled data layer built on Connect, available as a managed service today. Propexo sets up a Snowflake warehouse in an account that belongs to you, lands every system you run in it, and manages the whole thing end to end, with nothing for your team to build or host. Because the warehouse is yours, you decide which vendors and consultants get read-only access, through the warehouse’s own sharing controls, and you can revoke a grant any time. It is also what makes the intelligence layer buildable at all: every definition we fix and every identity we resolve has to hold against data the operator controls, and the managed warehouse is that ground. The semantics, the ontology, and the entity resolution this piece has been about are what we are building on top of it, and none of that machinery ships today.

Ordering is what the market keeps getting backwards. You cannot build a semantic layer over data you do not control. Every definition it fixes and every identity it resolves has to hold against systems that change their APIs on their own schedule, which means somebody has to own the movement layer underneath it permanently. A vendor who sells you the intelligence and rents the pipe from your PMS has built a roof with no walls, and the first vendor API change is the weather. Fivetran, Airbyte, and Matillion cover mainstream SaaS and databases well, and the multifamily operational stack is the layer they do not cover. Connect was built for exactly that layer, which is why the boring half came first and why it is the half becoming available today.

What to do about it before you buy another AI tool

Start by auditing one number. Pick a metric your team reports every month that crosses two vendors, and trace it back to which system’s definition of the shared field won. That exercise tells you more about your AI readiness than any vendor assessment or maturity model, and it takes an afternoon rather than a quarter. If the trace comes back clean, you are further along than most portfolios your size. And if it does not, you have just found the work.

Then get the data you own into a warehouse you own, whether your team runs it or we do, starting with the system whose absence has cost you the most. Not because a warehouse is the destination, but because every layer worth building sits on top of one, and the operators who will get real value from AI in eighteen months are the ones who spent this year on the boring part. The two customers we worked with most closely this year both had a warehouse before they had us. One was already paying for a horizontal platform and replaced the custom pipelines their own team had been maintaining into Snowflake. The other bought against a single named number, tour abandonment, and landed vendor data into a BigQuery warehouse they already ran. Neither needed convincing that data mattered, and neither engagement was frictionless. That pattern is not a coincidence, and it is the clearest signal we have about who is ready.

The industry spent three years learning how to move property data. Its next three go on learning what that data means. That second problem is harder, less demonstrable, and considerably more valuable, and the operators who take it seriously now will be the ones whose AI works when it matters.

Frequently asked questions

What changed for Propexo Connect today?
Connect is now available: open to any operator, no pilot gate, no waitlist. It has been running in production with paying customers for months, so this changes how you buy it rather than what it does. Support is included with every subscription through email and a shared Slack channel, onboarding runs two to six weeks depending on data readiness and source count, and Propexo holds SOC 2 Type II.
What is Connect Intelligence?
An owner-controlled data layer built on Connect, announced today and available as a managed service now. Propexo sets up a Snowflake warehouse in an account that belongs to you, lands every system you run in it, and manages it end to end, with partner access you grant and revoke through the warehouse sharing controls you own. The semantic layer, ontology, and entity resolution the name points toward are what Propexo is building on top of it, and none of that machinery ships today.
Why does AI need a semantic layer, an ontology, and entity resolution?
An agent is only as trustworthy as its grasp of what a word means, how things relate, and which records describe the same thing. Ask it to flag at-risk renewals and it has to know which records describe the same resident, how a lease relates to a unit and a unit to a property, and what at risk means in terms the data can express. Skip any of the three and it still answers, fluently and confidently, and the answer is wrong in a way nobody catches for a quarter.
Remen Okoruwa, Co-Founder & CEO, Propexo

Written by

Remen Okoruwa

Co-Founder & CEO, Propexo

Remen is co-founder and CEO of Propexo. A former McKinsey consultant and HubSpot Senior PM, he is a Harvard graduate and has passed all three levels of the CFA exam. He writes about the data infrastructure layer multifamily operators need before analytics or AI projects can ship.

Ready to get your property data where your team needs it?