Why modern matchmaking fails at 4 percent — and how applied psychology can do better.
Something like 350 million people worldwide have a dating app open on their phone right now, and almost none of them will marry the person they are currently swiping toward. This is not a cynical joke. It is close to the literal outcome of the modern dating market, and the industry that built it has made peace with the number.
Start with what we actually know, because the precise figures are contested and the apps that could resolve the argument mostly keep their data private. Pew Research Center's 2023 survey found that only about one in ten partnered U.S. adults met their current spouse or significant other on a dating site or app, rising to roughly one in five among partnered adults under 30 (Pew Research Center, 2023). Stanford sociologists Michael Rosenfeld, Reuben Thomas, and Sonia Hausen showed the more important structural shift underneath that number: meeting online overtook meeting through friends as the most common way American heterosexual couples get together sometime around 2013, and it has not looked back since, displacing not just bars and workplaces but the informal social networks that used to do the introducing (Rosenfeld, Thomas, & Hausen, 2019).
So dating apps did win. They became the dominant channel through which people meet partners. What they have not done is convert that dominance into proportionally more, or better, marriages. A recent NBER working paper by Ershov, Fong, and Yildirim, using two decades of U.S. county-level data on online dating usage from both the desktop era (2002–2013) and the mobile era (2017–2023), found that the mobile-era effect on marriage is negative: a 1 percent increase in online dating activity in a county is associated with a 0.40 percent decrease in the marriage rate and a 0.33 percent decrease in the divorce rate, meaning fewer marriages are forming in the first place rather than more failing ones (Ershov, Fong, & Yildirim, 2026). Independent work by Sunmi Jung and Lester Lusher, using variation in Tinder penetration across U.S. metropolitan areas as a natural experiment, reached a compatible conclusion: they find no evidence that dating apps have shifted the overall trajectory of marriage rates, with some evidence of increases in divorces among younger cohorts, and interpret the pattern as dating apps easing matching frictions in already-thick markets rather than manufacturing more lasting matches (Jung & Lusher, 2026). None of this means online dating cannot work. Roughly a fifth to a quarter of new marriages start there, which is a real number of real relationships. It means the conversion rate from match to marriage is bad enough that if a car had this failure rate, it would be recalled.
The scale involved makes the ratio worse than the headlines suggest. Tinder has reported, across various public disclosures, in the range of tens of billions of total matches since 2012 and roughly 1.5 million dates set up per week at various points in its history (Business of Apps, 2026). Even taking a generous view of how many matches convert to dates, and how many dates convert to relationships, and how many relationships convert to marriage, the number of marriages the platform can plausibly claim credit for each year is a rounding error against the volume of swipes that produced it. eHarmony, whose entire model was built around compatibility rather than browsing, has for years claimed responsibility for roughly 4 percent of U.S. marriages at its peak influence, a figure originating in company-commissioned Harris Interactive survey data and widely reported in the press (CNBC, 2015). That number is genuinely large in absolute terms and genuinely small as a fraction of the traffic it takes to produce it. Four percent is, roughly, the best a dating platform gets to claim as its share of a national outcome, after building one of the largest matching data sets ever assembled about human romantic behavior. That is the ceiling, not the floor, of the industry's ambition, and it is where this whitepaper gets its framing.
Why does an industry with this much data, this much engineering talent, and this much capital accept a conversion rate this low as normal? Three reasons, and none of them are secret.
The first is that the apps are not paid to produce marriages. They are paid, directly or through advertising, to produce engagement, and a person happily married is a person who deletes the app. Eli Finkel and his coauthors made this point plainly in their landmark 2012 review of the online dating industry: dating sites have no commercial incentive to solve the problem they claim to solve, because solving it ends the subscription (Finkel, Eastwick, Karney, Reis, & Sprecher, 2012). Hinge's marketing slogan, "designed to be deleted," is an admission dressed as a virtue: it concedes that the default mode of the category is not designed to be deleted, and treats leaving as an exception worth advertising rather than the expected outcome of using the product correctly.
The second reason is that the apps mostly collect the wrong inputs and process them the wrong way. Most matching happens on the same handful of signals: age, height, a few stated preferences, a handful of photos, and a feed algorithm that is, at bottom, optimized for the same engagement metrics as every other feed on your phone. Finkel and colleagues reviewed the psychological science behind mate matching algorithms in detail and concluded that the mathematical models used by most dating sites cannot work as advertised, because they are typically built on unvalidated proprietary formulas applied to superficial self-reported traits, rather than on the psychological constructs that the peer-reviewed literature actually shows predict relationship success (Finkel et al., 2012). More recent work by Eastwick, Finkel, and Joel formalizes this gap in "mate evaluation theory," building machine learning models that try to predict romantic interest from early relationship data and finding that even well-specified models struggle to outperform simple heuristics when they lean on the kind of static, stated preference data dating apps rely on (Eastwick, Finkel, & Joel, 2023; Joel, Eastwick, Allison, et al. dataset work). The tools are new. The inputs are not good enough to justify the tools.
The third reason is more structural than any single company's incentives: an oversaturated market with low switching costs rewards apps that keep you searching, not apps that get you a good match quickly. The Jung and Lusher result cited above is exactly this pattern in the data: more choice, faster search, cheaper rejection, without a corresponding improvement in the outcome the category claims to produce. None of that is the same thing as better matching.
We think the industry has quietly normalized a bad outcome because the outcome is not the product. Attention is the product. Marlow starts from the opposite premise: the outcome is the only thing that matters, and everything about the method should be built backward from it. That means asking a harder question than "what makes someone swipe right." It means asking what actually predicts whether two people will build something that lasts, and then building the entire system, uncomfortably, around the answer. The rest of this document is our attempt to answer that question honestly, and to show our work.
Ask someone what they're looking for in a partner and you will get a fluent, confident, well-organized answer. Ask a hundred people and you will get the same handful of answers in different words: kind, funny, ambitious, someone who "gets" them. This is not because everyone secretly wants the same person. It is because self-report, the method nearly every dating app relies on almost exclusively, is one of the weakest instruments in the entire behavioral sciences toolkit for predicting what a person will actually do.
The trouble is not that people lie about themselves, although some do. The deeper trouble is that people do not have reliable introspective access to the processes that actually drive their behavior, so even a scrupulously honest self-report is often just wrong. Daniel Kahneman's account of what he called WYSIATI, "what you see is all there is," describes a mind that builds a coherent, confident story out of whatever information happens to be available to it, without checking whether that information is representative or complete (Kahneman, 2011). Applied to dating profiles, this means a person describing their "ideal partner" is not consulting some buried, accurate ledger of their true preferences. They are generating a plausible story on the spot, built mostly from cultural scripts about what a good partner is supposed to sound like, then reporting that story with total sincerity.
This gap between what people say they want and what they actually choose has a name in economics: the difference between stated preference and revealed preference. Stated preference is what a person tells you in a survey. Revealed preference is what a person's actual choices show, when nobody is asking them to narrate their reasoning. The two routinely diverge, and not by a little. Eastwick and Finkel's own experimental work on ideal partner preferences found that people's stated ideals, the traits they claim matter most on a questionnaire, are strikingly poor predictors of who they actually pursue and become attracted to once they are interacting with real potential partners in person, as opposed to evaluating hypothetical profiles in the abstract (Eastwick, Finkel, & Eagly, 2011). Stated preferences do reasonably well at predicting reactions to a profile on a screen. They do much worse at predicting chemistry in a room. Since the entire dating app category is built around screen-based decisions, it has essentially optimized itself to be very good at the part of attraction that generalizes least to actual relationships.
This is not a new discovery, and it is not specific to romance. Consumer research has shown for decades that stated purchase intentions predict actual purchases only loosely. Political scientists know that stated voting intentions and issue priorities diverge from revealed behavior at the ballot box. The pattern recurs anywhere self-report is used to forecast a complex, socially loaded decision, because self-report captures the version of ourselves we present, consciously or not, rather than the mechanisms that generate our choices. Personality psychology has arrived at a related conclusion through a different door: people are decent judges of their own broad tendencies, like whether they're generally outgoing, but poor judges of how those tendencies will interact with a specific person in a specific context, which is exactly the prediction a matchmaking system needs to make.
There is also a more basic problem with asking better questions: the questions people expect on a dating profile have become so standardized that the answers have lost their signal. "What are you looking for in a partner." "What are your core values." These prompts now function more like a genre convention than an actual assessment. Everyone has learned the acceptable register of the answer, the same way everyone has learned to answer "tell me about yourself" in a job interview with a story about being a team player. The question stops measuring the person and starts measuring their familiarity with the format.
This is the argument this whitepaper is built around, so we will state it plainly: you cannot fix a self-report problem by writing a better self-report question. A cleverer prompt still asks the person to introspect and narrate, and introspection and narration are exactly the parts of the pipeline that the research shows are unreliable. The fix has to change what kind of evidence gets collected in the first place, shifting weight away from what a person says about themselves and toward what a person's patterns of behavior actually show. That does not mean discarding self-report. Validated psychometric instruments, administered properly, capture real and stable signal about personality and attachment, and Marlow uses several of them (see Chapter 3). It means self-report should never be the only signal, and it should be checked against behavior wherever behavior is available and consented to. Chapter 4 explains where we borrow that idea from, and Chapter 5 explains how we apply it without overreaching into the parts of a person's life that a matchmaker has no business touching.
If self-report alone is not enough, the obvious next question is: what does the research actually show predicts whether a relationship works. The honest answer is that decades of relationship science converge on a smaller and more specific set of constructs than the dating app industry currently uses. Marlow's matching method rests on four of them: Big Five personality traits, adult attachment style, values congruence, and conflict style. Each has a distinct research base, measures a different layer of how two people will actually live together, and has been shown, in peer-reviewed work, to relate to relationship outcomes. None of them, on its own, is a complete theory of love. Together, they are the closest thing the field has to a validated foundation.
Big Five personality. The Big Five model, sometimes called the Five-Factor Model, organizes personality into openness, conscientiousness, extraversion, agreeableness, and neuroticism (sometimes framed inversely as emotional stability). It is the most replicated structure in personality psychology, and it matters for matching for a specific reason: it is stable. Brent Roberts and colleagues, synthesizing decades of longitudinal data on how personality traits change across the lifespan, found that adult personality traits show strong rank-order stability over time, meaning that the relative ordering of people on these traits (who is more conscientious than whom) tends to persist for years, even as absolute trait levels shift gradually with age (Roberts, Walton, & Viechtbauer, 2006). A trait that is stable is a trait worth matching on, because it describes something durable about how a person is likely to behave next year and the year after, not just how they happen to feel on the day they filled out a questionnaire. A 2010 meta-analysis by John Malouff and colleagues, aggregating studies across thousands of couples, found that specific Big Five profiles reliably predict relationship satisfaction: higher agreeableness, conscientiousness, and emotional stability (low neuroticism) in a partner are consistently associated with a partner's own relationship satisfaction, while high neuroticism is one of the most consistent predictors of dissatisfaction and eventual dissolution across the studies reviewed (Malouff, Thorsteinsson, Schutte, Bhullar, & Rooke, 2010). This is not a small or tentative finding. It is one of the more robust results in relationship science, and almost no mainstream dating app measures it directly.
Adult attachment style. Attachment theory began as a framework for understanding how infants bond with caregivers, and psychologists Cindy Hazan and Phillip Shaver were the first to show, in a 1987 study that reshaped the field, that the same three attachment patterns identified in infants (secure, anxious, and avoidant) map cleanly onto how adults experience romantic love, including how they handle intimacy, jealousy, and separation (Hazan & Shaver, 1987). Mario Mikulincer and Phillip Shaver later built out the full architecture of adult attachment in relationships, showing across dozens of studies that a person's attachment orientation shapes how they regulate emotion during conflict, how much closeness they seek or avoid, and how they interpret a partner's behavior under stress (Mikulincer & Shaver, 2007). Two people can both be kind, both be funny, and still be a poor match if one has an anxious attachment style that intensifies under a partner's need for space, and the other has an avoidant style that creates exactly that space when stressed. This is not a personality clash in the Big Five sense. It is a mismatch in how two nervous systems respond to closeness, and it is one of the more predictable sources of recurring conflict in long relationships.
Values congruence. Social psychologist Shalom Schwartz developed a widely validated model of universal human values, organizing the things people prioritize (achievement, security, benevolence, tradition, and others) into a structured circle where some values naturally support each other and others are in tension (Schwartz, 1992). Values congruence between partners, the degree to which two people prioritize similar things even if they express that in different personalities, has been linked in subsequent relationship research to relationship satisfaction and longevity, largely because values shape the thousands of small decisions a couple makes about money, time, family, and risk long after the initial attraction has settled into routine. Two agreeable, secure, emotionally stable people can still be a poor match if one organizes their life around stability and tradition and the other around novelty and self-direction, because those decisions compound over a shared life in a way that a single good date cannot reveal.
Gottman conflict style. John Gottman's decades of observational research on married couples, conducted with longitudinal follow-up rather than one-time surveys, identified specific communication patterns during conflict that predict marital dissolution with notable accuracy. Gottman and Nan Silver's synthesis of this research names four corrosive patterns, criticism, contempt, defensiveness, and stonewalling, that Gottman's lab nicknamed the Four Horsemen, and showed that the presence of contempt in particular is one of the single strongest predictors of divorce in his longitudinal samples (Gottman & Silver, 1999). This matters for matching because conflict style is not really about whether two people fight. Every real relationship has friction. It is about how two people fight, and whether their particular styles escalate or de-escalate each other. A person who withdraws under stress paired with a person who needs to talk through conflict immediately can produce a stable, repeating cycle of frustration that neither person's individual personality would predict on its own.
Each of these four frameworks captures something the others miss. Big Five is about temperament: the steady traits a person brings into any room. Attachment is about intimacy regulation: what happens to that temperament under the specific pressure of closeness and separation. Values are about direction: what a life organized around this person will actually prioritize. Conflict style is about repair: what happens when direction and temperament collide, as they eventually do in every relationship. A matching system that only looks at one of these is measuring a slice of the person and calling it the whole picture. Marlow measures all four, and Chapter 8 details exactly which validated instruments we use to do it.
Investigative psychology, the discipline that produced modern behavioral analysis in criminal investigation, was built on a principle that long predates any particular unit or agency: watch what someone does, not what they say they'll do. An investigator reconstructing a person from evidence is not asking that person to fill out a self-report survey about their personality. They are inferring psychology from behavior, because behavior, especially patterned, repeated behavior, is harder to fake than a single statement and harder to spin than a self-description.
Two specific tools from that tradition are worth borrowing carefully, and worth naming precisely, because Marlow's method leans on the underlying logic without borrowing the investigative framing wholesale.
The first is baseline and deviation analysis. Investigators trained in behavioral analysis do not look at a single gesture, statement, or reaction in isolation and declare it meaningful. They first establish what is normal for a specific individual, their baseline, through low-stakes observation, and only then look for deviations from that baseline when a sensitive topic arises. Paul Ekman, whose research on nonverbal behavior heavily influenced this approach, described the method directly: establish a baseline of a person's ordinary behavior, then watch for departures from it, because the departure is informative in a way that no universal gesture is (Ekman, cited in Paul Ekman Group, 2021). This is a meaningfully more modest claim than the pop-culture version of "reading people," and it is worth being precise about the limits here, because the science is more cautious than the mythology. A substantial body of subsequent research, including a widely cited review by Aldert Vrij and colleagues, has found that most individual nonverbal cues popularly associated with lying (gaze aversion, fidgeting, pauses) are, on their own, faint and unreliable predictors of deception, and that verbal content is often more diagnostic than body language (Vrij, Hartwig, & Granhag, 2019; DePaulo et al., 2003). We take that caution seriously rather than ignoring it. The lesson Marlow draws from this literature is not "we can detect lies from body language." It is the narrower, better-supported idea underneath it: a single data point about a person, viewed with no reference point, tells you very little. The same data point, viewed against that specific person's own pattern, tells you much more. Baseline-relative analysis, not universal cue-reading, is the part of the tradition that has held up.
The second tool is statement analysis, sometimes called content-based analysis, a family of techniques developed within investigative psychology for evaluating the structure and content of a person's account of events, independent of the person's demeanor while giving it. The core insight, refined across decades of applied investigative work, is that what people spontaneously choose to include, omit, elaborate on, or avoid in their own account of their life reveals more than a direct question ever could, precisely because the person is not aware they are being assessed on that dimension. A person's Spotify listening pattern, unprompted and unperformed, is a mundane version of the same idea: nobody curates their commute playlist for a dating profile.
This connects directly to the argument in Chapter 2. The reason self-report is weak is that it is a performance, however sincere, given in the knowledge that it will be judged. Behavior collected with consent but without a performance context, an Instagram social graph, a Spotify listening history, is not a performance in the same way. A person choosing who to follow on Instagram over years is not thinking "how will a matchmaker interpret this," the way they are when writing a dating bio. That absence of performance pressure is exactly what makes the behavioral signal more informative than the self-report, and it is exactly the same shift in evidentiary standard that took criminal profiling from asking suspects what kind of person they are to reconstructing what kind of person the evidence shows them to be.
We want to be direct about where the analogy ends. An investigator is trying to identify a specific unknown subject from limited, often adversarial evidence, working under enormous asymmetry between what the subject knows and what the investigator knows. Marlow is trying to build an accurate, humane picture of a consenting member who wants to be understood, with full visibility into what data is being used and full ability to withdraw it. The stakes, the consent structure, and the goal are entirely different, and we are not in the business of psychological surveillance. What we take from investigative psychology is not its subject matter. It is a single methodological commitment: behavior observed without a performance incentive is a more reliable window into a person's actual patterns than a description that person gives of themselves under the pressure of being judged. This commitment is echoed in a parallel line of research on digital behavior, where Kosinski and colleagues showed that ordinary online traces (Facebook likes, in their study) can predict Big Five personality traits with accuracy comparable to close friends and family, and Youyou and colleagues later showed that computer-based judgments from those traces can outperform human judgments of personality made by acquaintances (Kosinski, Stillwell, & Graepel, 2013; Youyou, Kosinski, & Stillwell, 2015). The direction of that evidence is the same as ours: unperformed behavior, taken at sufficient scale, carries real personality signal. Chapter 5 explains exactly which behavioral sources we use, how narrowly we scope them, and Chapter 6 draws the boundary around everything we refuse to touch even though the industry increasingly has access to it.
Everything before this chapter is argument. This chapter is architecture. It describes, at the level a curious member or a critical researcher would want, how Marlow actually turns the four frameworks in Chapter 3 and the behavioral evidence in Chapter 4 into a single weekly introduction between two specific people. Marlow's method is delivered through Aria, the matchmaker persona members interact with throughout the experience, whose role is to make an otherwise cold assembly of scores and signals feel like a human handoff. Aria is a design choice, not a claim of intelligence, and we come back to what she is and is not in Chapter 7.
The pipeline has four stages: inputs, profile, matching, introduction.
Stage 1 — Inputs. Three streams feed a member's profile, each doing work the other two cannot.
Personality and relational assessments, covering the four frameworks in Chapter 3: the Mini-IPIP for Big Five, the ECR-R short form for attachment, a Schwartz-derived values card sort, and a set of Marlow-original conflict scenarios inspired by Gottman's observational work. Full instrument detail is in Chapter 8. These instruments carry the bulk of the psychological signal.
Consented behavioral data, used as a check and a supplement, never as a replacement for the assessments. With explicit, revocable consent, a member connects Instagram (analyzed as a social graph — the shape and composition of who they follow, not image content) and Spotify (listening patterns over time via OAuth). Facebook and Goodreads are optional for members who want a richer profile. This layer exists because of the argument in Chapters 2 and 4: unperformed behavior is a more honest signal than a self-description written for an audience.
Life-stage questions, a short set of direct questions about timeline, family goals, location flexibility, and relationship structure. These are the questions self-report is actually good at answering, because they concern facts and near-term intentions rather than deeper patterns.
Stage 2 — Profile. The three input streams are compiled into a single structured member profile with a deliberately small number of dimensions: five Big Five trait scores, two attachment dimensions (anxiety and avoidance), a rank-ordered values vector, a conflict-style classification, and a life-stage record. The behavioral data does not add new dimensions. It is used two ways. First, as a coherence check: where the Instagram and Spotify patterns diverge sharply from the assessment answers on the dimensions where behavior has been shown to carry signal (openness and extraversion in particular), the divergence is flagged for review before a match is generated, on the premise from Chapter 2 that a person's actual patterns should be given weight when they contradict a self-description. Second, as color rather than score: behavioral texture (the fact that a member's listening skews heavily toward late-night acoustic music, or that their follow graph clusters around independent bookstores) informs how Aria writes the Reading and the introduction, not the compatibility score itself.
The finished profile is returned to the member as a personalized Reading — a plain-language reflection of who the data suggests they are, written in Aria's voice, warm rather than clinical. The Reading is not decorative. It is the mechanism by which the member can recognize, correct, or push back on the picture Marlow has formed of them before it is ever used to generate a match. A member should never feel matched by an algorithm they cannot see the outline of, and the Reading is the outline.
Stage 3 — Matching. Life-stage acts as a hard filter, not a weight. Two members with incompatible timelines or non-overlapping cities are never scored against each other, because no amount of personality compatibility redeems a fundamental mismatch on what each person is looking for. Within the filtered candidate pool, compatibility is scored across the four frameworks jointly. Two properties of this scoring are worth naming clearly, because they are where most matching systems quietly fail.
The first is that different frameworks reward different relationships between two members. Values reward congruence: two people are a better match when they prioritize similar things, because values compound across a shared life. Big Five is more nuanced: agreeableness, conscientiousness, and emotional stability reward high absolute levels in both partners (the Malouff meta-analysis in Chapter 3 is unambiguous on this), while openness and extraversion tolerate wider gaps and sometimes benefit from moderate difference. Attachment rewards a specific structural fit — two secure partners is the strongest configuration, and certain insecure pairings (anxious-avoidant in particular) are structurally destabilizing regardless of how well matched the pair looks on other dimensions. Conflict style rewards complementarity in repair capacity: two people whose styles de-escalate rather than escalate each other, which is not the same as two people with the same style.
The second is that the frameworks are scored jointly, not sequentially. A high score on three frameworks does not compensate for a structural incompatibility on the fourth. In practice this means the engine applies a minimum threshold on each of the four dimensions before a joint score is even computed, so that a pair strong on values, personality, and conflict style but structurally mismatched on attachment does not surface as a top match on the strength of the other three. The joint score is used only to rank candidates who have already cleared every framework's floor. Specific weights and thresholds are proprietary, for the same reason a locksmith does not publish a master key, but every input feeding this scoring is disclosed here and in Chapter 8, and nothing enters the engine that is not named in this document.
The engine surfaces one candidate per week, delivered on Sunday. It does not surface a feed. The pacing is a deliberate consequence of Chapter 1's argument: an endless queue optimizes for engagement, and one considered introduction optimizes for outcome.
Stage 4 — Introduction. When a match clears, Aria writes both members a short, specific introduction to the person they are about to meet: who they are, and why Marlow believes the match makes sense, grounded in the actual profile rather than generic flattery. The introduction is the point at which the abstract scoring becomes a human sentence, and it is what a member actually experiences of the method.
After a match happens, Marlow's work is not done. We track outcomes longitudinally at 6, 12, and 24 months post-match, asking members what happened and how the relationship, if one formed, is going. This exists for two reasons. It lets Marlow's matching engine improve against a real outcome measure instead of a proxy like messages sent or dates scheduled. And it is the mechanism by which we intend to eventually publish real, dated, first-party data on our own match-to-relationship and match-to-marriage rates, rather than asking members to trust an unverifiable claim. Chapter 7 addresses directly why we do not have that data yet, and Chapter 8 details the tracking protocol.
Trust in a system like this cannot rest on a promise to be careful. It has to rest on a public, specific list of the data Marlow refuses to collect, published in full, with the reasoning behind each refusal, so that the commitment can be checked rather than taken on faith. This is that list.
Search history. We will never access what a member searches for, on any platform. Search history is arguably the single most intimate record a person generates, more revealing of fear, desire, and private struggle than almost anything else they produce, precisely because it is composed in the belief that no one is watching. Using it to build a romantic profile would mean quietly weaponizing the one place a person expects total privacy, and we consider that expectation worth protecting absolutely rather than worth exploiting for match quality.
Text messages and emails. We will never read a member's private correspondence. A conversation between two people belongs to both of them, and neither one consented to a matchmaking algorithm reading it, even if one of them consented to Marlow generally. Group and private conversations carry other people's privacy inside them that a member has no right to sign away on their behalf.
Purchase history. We will never access what a member buys. Purchase data is a favorite input for advertising systems because it reveals income, health conditions, family status, and vice with startling precision, and that precision is exactly the problem: it would let Marlow infer things about a member's finances, medical situation, or private habits that have nothing to do with romantic compatibility and everything to do with surveillance capacity.
Location beyond city. We will know what city a member lives in, because location within a reasonable radius is a basic practical requirement of matchmaking. We will never track precise or continuous location. Fine-grained location data can reveal where someone works, worships, seeks medical care, or spends nights, and none of that belongs in a dating profile.
Health data. We will never access medical records, fitness tracker biometrics, mental health app data, or insurance information. Health status can correlate with almost anything, including things a person has every right to keep private from a matchmaking company, and using it would blur the line between "helping you find a partner" and "assessing your body," a line we do not intend to approach.
Financial data. We will never access bank accounts, credit scores, income, or debt. Financial compatibility is a real and valid thing for two people to discuss with each other directly, but it is not Marlow's business to score or gate a person's romantic prospects by their bank balance, and doing so would import a class filter into a matching system that has no place there.
Photo content. We will use aggregate expression signal where relevant to consented profile-building (for example, whether Instagram photos as a set skew toward outdoor activity or social gatherings), but we will never analyze the content of individual images, run facial recognition, or make judgments based on appearance. Attractiveness is real and Marlow does not pretend otherwise, but automated appearance scoring imports exactly the shallow, first-glance judgment that this whitepaper argues is already overrepresented in dating apps, and it is not a signal we will systematize.
Data brokers. We will never buy data about a member from a third party. Every piece of information Marlow holds about a member will have been given directly, knowingly, and revocably by that member. A data broker relationship would sever that chain of consent entirely, and consent is the foundation the whole method stands on, not a compliance checkbox.
Streaming viewing history. We will never access what a member watches on Netflix or similar services. Viewing history sits close to search history in intimacy, often reflecting private interests, identity exploration, or comfort-seeking that a person has not chosen to share, and it offers little compatibility signal that isn't better captured by the consented sources we do use.
We are aware that every item on this list is something a sufficiently motivated technology company could obtain, and that some competitors already infer from adjacent signals what they cannot access directly. The list exists precisely because "we could" is not the same question as "we should." It is a boundary we are choosing to hold even when holding it costs us predictive accuracy we could otherwise buy.
A method that only lists its strengths is not being fully honest with the people relying on it. This chapter is where we say, plainly, where Marlow's approach is unproven, untested, or genuinely uncertain, because a claim to scientific credibility that omits its own limitations is not actually a scientific claim.
We do not yet have our own long-term outcome data. Everything in Chapter 3 is drawn from decades of published relationship science, and it is real evidence that these four frameworks predict outcomes in general populations. It is not yet evidence that Marlow's specific implementation of these frameworks, combined with consented behavioral data and Aria's matchmaking layer, produces better outcomes than the industry status quo. That evidence can only come from Marlow's own longitudinal tracking, described in Chapter 5, and it takes time to accumulate honestly. We are committing publicly to reporting our own 24-month outcome data once we have a cohort large enough to report on responsibly, whatever that data shows.
Our frameworks carry a Western-centric bias. The Big Five, attachment theory, and Schwartz's values model were developed primarily on Western, educated, and disproportionately American and European samples, and cross-cultural psychology has repeatedly found that personality structures, attachment norms, and value hierarchies vary in ways that are not fully captured by instruments built on those samples. The Mini-IPIP, ECR-R, and Schwartz-derived instruments Marlow uses are well validated within the populations they were built and tested on. We have not yet validated them separately across the full range of cultural backgrounds present in Marlow's membership, and we should not overstate their universality until we have.
We do not know what effect Aria herself has on outcomes. Aria's Readings and introductions are designed to make the process warmer and more legible, but an AI matchmaker character is itself an intervention, not a neutral delivery mechanism, and we do not yet know whether it changes how members show up to a match, how they interpret compatibility information, or whether they trust a match more or less than they would trust the same information delivered without a character attached to it. This is a testable question and we intend to test it, but we have not yet done so with the rigor the question deserves.
Marlow's membership is self-selected, and that changes what any result would mean. People who choose a $49-a-month matchmaking service with a lengthy assessment process are not a random sample of people looking for a partner. They are likely more relationship-serious, more patient with process, and probably wealthier than the average dating app user, on average. Any outcome data Marlow eventually publishes will describe what happens for people like our members, not what would happen if this method were applied to the dating population at large, and we will be explicit about that boundary when we report results rather than letting readers assume otherwise.
We think this chapter is the most important one in the document, not the least important. A method this specific, applied to something as consequential as a person's romantic life, deserves scrutiny in proportion to its confidence. We would rather a skeptical reader trust us more after this chapter than trust us blindly before it.
This appendix describes the specific instruments and technical protocols behind the method summarized in Chapter 5, at the level of detail a researcher or journalist would need to evaluate the approach. We do not publish the proprietary weighting or architecture of the matching engine itself, for the same reason a locksmith does not publish a master key, but every input, instrument, and outcome protocol is described here in full.
Big Five instrument. Marlow uses an adapted Mini-IPIP, a 20-item short form of the International Personality Item Pool developed and validated by Donnellan, Oswald, Baird, and Lucas (2006) as a compact, psychometrically sound alternative to longer Big Five inventories, well suited to a mobile onboarding context without sacrificing the scale's internal reliability. Marlow presents select items as forced-choice pairs adapted from the original Likert format, a design choice intended to reduce the flattering, midpoint-clustering response bias common in self-report personality scales, while preserving the underlying five-factor structure the instrument is validated to measure.
Attachment instrument. Marlow uses the ECR-R short form, a 12-item version of the Experiences in Close Relationships scale validated by Wei, Russell, Mallinckrodt, and Vogel (2007), which measures attachment along the two empirically supported dimensions of anxiety and avoidance rather than forcing members into a single categorical attachment label. This is the same underlying construct structure established by Hazan and Shaver (1987) and elaborated by Mikulincer and Shaver (2007), operationalized in the abbreviated form the Wei et al. validation work established as reliable for use outside long research batteries.
Values instrument. Marlow's values assessment is a card-sort exercise adapted from Schwartz's Portrait Values Questionnaire (Schwartz, 1992), asking members to rank and prioritize value statements against each other rather than rate them independently, which produces a more accurate relative priority structure than independent Likert ratings tend to yield, since values are, by nature, expressed in tradeoffs against each other rather than in isolation.
Conflict scenarios. Marlow presents members with short, concrete conflict scenarios (a canceled plan, a financial disagreement, a disagreement about how to spend a holiday) and asks how the member would respond in the moment. These scenarios are custom-developed for Marlow, inspired by the communication patterns Gottman identified in his observational marriage research (Gottman & Silver, 1999), but they are not a licensed or standardized Gottman instrument, and we describe them here as Marlow-original so that this distinction is never ambiguous to a researcher evaluating our claims.
Behavioral data sources. Instagram data is collected as a social graph (following and follower relationships and account-level metadata) via the member's authenticated handle, with explicit scoped consent and no access to private message content. Spotify data is collected via OAuth with PKCE (Proof Key for Code Exchange), a secure authorization flow that lets a member grant Marlow read access to their listening history without ever sharing their Spotify password, and which the member can revoke at any time through their own Spotify account settings, immediately cutting Marlow's access. Facebook and Goodreads connections, where offered, follow the same consent and revocation standard.
Longitudinal tracking protocol. Matched members receive brief, optional outcome surveys at 6, 12, and 24 months following an introduction, asking whether the match progressed to a relationship and, where applicable, how it is going along basic satisfaction and stability measures drawn from established relationship satisfaction research. Responses are voluntary; non-response is tracked as its own category rather than silently excluded, so that published outcome rates are not inflated by selection bias among the members most eager to report good news.
Matching engine architecture, high level. Member profiles, once built from the assessments and consented behavioral data above, are represented across the four frameworks in Chapter 3 plus life-stage fit. The engine evaluates compatibility across all dimensions jointly rather than sequentially, ranks candidate matches for a given member above a minimum compatibility threshold, and surfaces one candidate per week. We do not publish the specific weighting formula or scoring architecture, consistent with standard practice for any applied matching system, but every input feeding that architecture is listed in full in this appendix and nowhere in the system does an input appear that is not disclosed somewhere in this document.
Business of Apps. (2026). Tinder Revenue and Usage Statistics (2026). https://www.businessofapps.com/data/tinder-statistics/
CNBC. (2015, May 8). How to surf the Web for a mate: eHarmony founder. https://www.cnbc.com/2015/05/08/how-to-surf-the-web-for-a-mate-eharmony-founder.html
Donnellan, M. B., Oswald, F. L., Baird, B. M., & Lucas, R. E. (2006). The Mini-IPIP scales: Tiny-yet-effective measures of the Big Five factors of personality. Psychological Assessment, 18(2), 192–203.
Eastwick, P. W., Finkel, E. J., & Eagly, A. H. (2011). When and why do ideal partner preferences affect the process of initiating and maintaining romantic relationships? Journal of Personality and Social Psychology, 101(5), 1012–1032.
Eastwick, P. W., Finkel, E. J., & Joel, S. (2023). Mate evaluation theory. Psychological Review, 130(1), 211–241.
Ershov, D., Fong, J., & Yildirim, P. (2026). What Happens When Dating Goes Online? Evidence from U.S. Marriage Markets and Health Outcomes (NBER Working Paper No. 34757). National Bureau of Economic Research. https://www.nber.org/papers/w34757
Finkel, E. J., Eastwick, P. W., Karney, B. R., Reis, H. T., & Sprecher, S. (2012). Online dating: A critical analysis from the perspective of psychological science. Psychological Science in the Public Interest, 13(1), 3–66.
Gottman, J. M., & Silver, N. (1999). The Seven Principles for Making Marriage Work. Crown Publishers.
Hazan, C., & Shaver, P. (1987). Romantic love conceptualized as an attachment process. Journal of Personality and Social Psychology, 52(3), 511–524.
Jung, S., & Lusher, L. (2026). Dating apps and marriage rates. Economics Letters, 262.
Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110(15), 5802–5805.
Malouff, J. M., Thorsteinsson, E. B., Schutte, N. S., Bhullar, N., & Rooke, S. E. (2010). The Five-Factor Model of personality and relationship satisfaction of intimate partners: A meta-analysis. Journal of Research in Personality, 44(1), 124–127.
Mikulincer, M., & Shaver, P. R. (2007). Attachment in Adulthood: Structure, Dynamics, and Change. Guilford Press.
Pew Research Center. (2023, February 2). Key findings about online dating in the U.S. https://www.pewresearch.org/short-reads/2023/02/02/key-findings-about-online-dating-in-the-u-s/
Roberts, B. W., Walton, K. E., & Viechtbauer, W. (2006). Patterns of mean-level change in personality traits across the life course: A meta-analysis of longitudinal studies. Psychological Bulletin, 132(1), 1–25.
Rosenfeld, M. J., Thomas, R. J., & Hausen, S. (2019). Disintermediating your friends: How online dating in the United States displaces other ways of meeting. Proceedings of the National Academy of Sciences, 116(36), 17753–17758.
Schwartz, S. H. (1992). Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries. Advances in Experimental Social Psychology, 25, 1–65.
Vrij, A., Hartwig, M., & Granhag, P. A. (2019). Reading lies: Nonverbal communication and deception. Annual Review of Psychology, 70, 295–317.
Wei, M., Russell, D. W., Mallinckrodt, B., & Vogel, D. L. (2007). The Experiences in Close Relationships Scale (ECR)-Short Form: Reliability, validity, and factor structure. Journal of Personality Assessment, 88(2), 187–204.
Youyou, W., Kosinski, M., & Stillwell, D. (2015). Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences, 112(4), 1036–1040.
— The Marlow research team