From Decks to Decisions: An Auditable AI Framework for Venture Capital | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article From Decks to Decisions: An Auditable AI Framework for Venture Capital Amin Haghani, Haydar Nur, Serhat Pala This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7781525/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Venture capital faces a paradox of scale and scrutiny: rising volumes of inbound deals, increasingly multimodal pitch materials, and higher expectations for rigor and fairness. We present an auditable AI framework that converts pitch decks into a rubric-aligned, evidence-linked embedding using a multimodal model, then trains leakage-aware classifiers for a progressive funnel (prescreen, pitching, due diligence). On held-out tests, the best stage models achieved AUCs of 0.717 (prescreen), 0.692 (pitching), and 0.836 (due diligence). In a prospective cohort (n = 281), high-confidence thresholds shaped workload as intended: 59.1% passed prescreen, 10.2% of those advanced to pitching (6.0% of total), and 0.0% advanced to due diligence, reflecting deliberate final-stage stringency. Feature profiles aligned with investor priorities: geography and team capacity at intake; market realism, team execution, and defensible differentiation at diligence; mid-funnel effects were smaller but interpretable. Unlike generic or opaque pipelines, our approach centers governance—direct citations, strict leakage controls, test-set model selection—turning thresholds into policy dials and AI into a documented collaborator. The methodology is transferable (re-elicitation required), and prospective validation with human–AI collaboration studies and active-learning loops is a practical next step for trustworthy deployment in private markets. Artificial Intelligence and Machine Learning Operations Research Finance Entrepreneurship Artificial Intelligence LLM Venture Capital Machine Learning Investment Screening Explainable AI Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Venture capital (VC) faces a paradox of scale and scrutiny: record volumes of inbound opportunities, increasingly complex materials (pitch decks, KPI tables, product roadmaps), and heightened expectations for rigor and fairness in selection. Manual, network-driven screening strains under this load and can be inconsistent. At the same time, recent advances in large language models (LLMs)—especially multimodal models capable of reading text, tables, and figures—have shown promise in document triage and synthesis when paired with disciplined evidence and governance practices (Guo et al. 2024 ; Yanglet, Cao, and Deng 2025). Most prior approaches to VC screening are either generic (optimizing against public proxies rather than a fund’s thesis) or opaque (limiting adoption in regulated contexts). Text-centric pipelines also miss the visual and structural content that conveys signal in pitch decks (e.g., traction charts, unit-economics tables, product diagrams). Methodological pitfalls—data leakage, look-ahead bias—further inflate performance estimates, and black-box systems inhibit trust where provenance and explainability are required (Vollmer et al. 2020 ). By contrast, rubric-aligned outputs—where the AI scores against predefined criteria with rationales—improve interpretability and reliability for downstream decisions and assessments (Askari 2025 ). We introduce an applied, auditable AI framework that addresses these gaps by aligning analysis and modeling to a fund-specific investment rubric. First, we elicit investor criteria through interviews and a structured questionnaire to encode a rubric that reflects the committee’s actual decision factors. Second, a multimodal model (Google Gemini 2.5 Pro) reads pitch decks and contextual materials and produces a structured embedding—criterion-level scores, rationales, and direct evidence links—enforcing a stable schema and reducing free-text variance (Guo et al. 2024 ; Yanglet, Cao, and Deng 2025). Third, we transform this embedding into features for a three-stage machine-learning funnel (prescreen, pitching, due diligence) with strict leakage controls and test-set model selection to ensure honest evaluation (Vollmer et al. 2020 ). Finally, we generate investor-facing HTML summaries with inline citations and print-ready PDFs under a governance regime emphasizing provenance, auditability, and reproducibility. Contributions are threefold. (i) A multimodal, rubric-aligned analysis layer that converts heterogeneous pitch materials into verifiable, structured signals. (ii) A leakage-aware ML pipeline tailored to VC’s progressive funnel, with stage-specific thresholds calibrated as operational policy levers. (iii) A governed path from raw decks to decision support that foregrounds explainability and evidence, suitable for peer review and operational deployment (Vollmer et al. 2020 ; Askari 2025 ; Guo et al. 2024 ; Yanglet, Cao, and Deng 2025). We position this as an exemplar implementation at NuFund; the methodology is transferable, though adoption elsewhere requires re-eliciting investor indicators and validating outputs against local decision criteria. Results Human Survey & AI Training Loop We implemented a continuous human feedback mechanism that enables ongoing refinement of the AI system based on investor evaluations. This component collects structured feedback from NuFund investors on AI-generated predictions and reports, allowing for rubric refinement, threshold calibration, and model improvement. The feedback loop ensures the AI system remains aligned with evolving investor preferences and market conditions, creating a dynamic learning system that improves over time. Investor Survey and Rubric Elicitation We conducted semi-structured interviews and a structured questionnaire to elicit decision criteria and operational cutoffs from NuFund investors. The final rubric distilled recurring indicators across market, team, product, traction, and financing terms into bounded, interpretable sub-scores and risk flags, which the AI layer uses to generate a structured, evidence-linked embedding for each deal (Fig. 1 ). This process standardized language across reviewers and enabled consistent application of stage-specific thresholds downstream. The initial survey insights now feed into the continuous Human Survey & AI Training loop for ongoing refinement. Key findings from interviews (synthesized): Founders and team Strong founder–market fit and complementary founding teams were universally emphasized; solo founders at POC are higher risk. Integrity and commitment checks (fraud/bankruptcy/legal history; time dedication) are mandatory screens; inflated salaries at early stage are a red flag. Market and competition Validated market research and credible SAM/TAM assumptions are required; over-broad or inflated markets are a red flag. Strategic small markets can be acceptable if there is clear strategic importance; otherwise small/poorly defined markets are disfavored. Product and roadmap Scalability and feasibility must be evidenced; over-customization and unrealistic roadmaps are negatives. Clear differentiation and awareness of direct/indirect competitors are expected. Financials and terms Realistic projections and clean cap tables are baseline; misleading metrics (e.g., bookings vs ARR) or heavy debt are red flags. Capital efficiency and a credible cash-flow/raise plan are positives; dependency on large unfunded raises is risky. Legal/regulatory and presentation Early regulatory strategy (e.g., FDA) and IP posture should be addressed in regulated sectors. Deck quality matters: concise, professional, error-free, with defensible numbers; cluttered/error-filled decks signal risk. Distinct emphases by investor Risk focusing (avoid taking high market and high product risk simultaneously); commercialization plans in med-device; diversity as a team strength; skepticism of AI over-claims without technical depth; acquisition potential in life sciences. These insights directly informed the 24-criterion rubric (− 10 to + 10 per item) spanning founders, market, product, financials, regulatory, and presentation, yielding a total score range of − 240 to + 240 with qualitative interpretation bands for screening (Fig. 1 ). Database and Funnel Characteristics Our Dealum cohort comprised 487 unique companies. Most submissions were early-stage: 57.7% at intake (Applied/Referral), while 15.8% were ultimately Funded and 16.4% Rejected/Archived. Industry composition (known for 80.7% of deals) was led by Healthcare/Biotech (44.0%) and Technology/AI (30.0%). Companies typically sought modest raises (median 1.28M USD; mean 2.13M USD; range 7.25K USD to 65.0M USD) and were concentrated at pre-seed/seed (76.0% of those reporting stage). Prior funding was sparse (median 0.75M USD; mean 2.34M USD). Figure 2 summarizes the funnel, top industries, and capital/stage/funding distributions. The progressive workflow comprises three stages—prescreen, pitching, and due diligence—each calibrated to shape workload and selectivity. On an independent prospective set of 281 new deals, high-confidence thresholds yielded: prescreen 59.1% pass, 10.2% of prescreen passes advancing to pitching (6.0% of total), and 0.0% advancing to due diligence (reflecting stringent final-stage selectivity). For operational practice, we also used relaxed decision thresholds (0.8, 0.6, 0.5 for prescreen, pitching, due diligence) to flag candidates for review while reserving 0.823/0.685/0.704 as high-confidence cutoffs. AI-Generated Embedding Analysis The AI system generates structured embeddings for each deal based on the 24-criterion rubric, creating 10-dimensional feature vectors that capture key assessment dimensions: Team Expertise, Market Opportunity, Business Model, Competitive Advantage, and Financial Projections. These embeddings are normalized to a 0–10 scale and provide quantitative representations of deal quality across multiple dimensions. Figure 3 shows the distribution of these AI-generated scores across funnel stages, revealing distinct patterns in how different stages cluster around specific score ranges. The correlation analysis demonstrates relationships between assessment dimensions, while the radar chart illustrates how mean scores vary systematically across funnel stages, highlighting the progression of deal quality through the investment process. Progressive ML Performance Models were selected by test-set AUC per stage. The best individual models were GradientBoosting (prescreen, AUC 0.717), SVC (pitching, AUC 0.692), and SVC (due diligence, AUC 0.836). Thresholds used operational cutoffs aligned to workload and selectivity (prescreen 0.823; pitching 0.685; due diligence 0.704 for high-confidence reporting). Stage-level ROC, PR, confusion matrix, and sensitivity/specificity panels are provided in Fig. 4 (A–D) and supplemental panels. Feature importance and feature set All stages use an identical 25-feature set (15 structured variables such as industry flags, stage and geography encodings, capital seeking, web presence; and 10 AI-derived rubric scores spanning team, market, product, finance, and scalability). Prescreen (GradientBoosting): The most influential features were country_encoded (24.2%), Team_Expertise (20.1%), capital_seeking_numeric (8.4%), Financial_Projections (8.3%), and company_stage_encoded (7.9%), followed by Competitive_Advantage and Market_Opportunity (~ 6–7%). This mirrors interview priorities that emphasized strong teams, realistic funding asks/projections, and stage fit. Pitching (SVC): Same feature set; the figure shows country_encoded (5.6%) as the top feature, followed by capital_seeking_numeric (1.4%), Market_Opportunity (0.6%), has_digital_health (0.3%), and has_website (0.3%); most other features show zero importance. This aligns with prescreen patterns where geographic and financial factors dominate early-stage decisions, though with much lower overall importance than the prescreen stage. Due diligence (SVC, permutation importance): Market_Opportunity was the top signal (~ 3.5%), followed by Team_Expertise (~ 3.2%) and Competitive_Advantage (~ 2.8%); sector and stage signals were smaller (has_digital_health ~ 2.5%, funding_stage_encoded ~ 2.2%). This aligns with interviews that final decisions emphasize market realism, team execution capacity, and defensible differentiation, with sector/stage context contributing but not dominating. Discussion Our findings demonstrate that a governed, rubric-aligned, multimodal AI layer can translate heterogeneous pitch materials into auditably structured signals that are both usable by downstream ML and verifiable by investors. In contrast to much of the startup-prediction literature that relies on generic public features and text-only embeddings—and is often weakened by look-ahead bias or leakage—our approach pairs an investor-elicited rubric with a multimodal model to capture visual and tabular signal from decks, then enforces verifiable provenance through inline citations (Guo et al. 2024 ; Vollmer et al. 2020 ; Yanglet, Cao, and Deng 2025). This design directly addresses two gaps highlighted in prior work and practice: the absence of fund-specific criteria in modeling pipelines and the opacity of “black-box” systems that hinder adoption in regulated contexts. By forcing the AI to produce rubric-aligned scores, rationales, and evidence links, we shift from free-text outputs to an interpretable schema that improves both machine learning stability and human review, consistent with evidence that rubric-aligned assessments enhance transparency and reliability (Askari 2025 ). Empirically, three stage-wise patterns connect investor judgment with model behavior. First, at prescreen, geographic context and team capacity dominate alongside pragmatic financing signals; this mirrors interview themes around founder-market fit, execution readiness, and credible capital plans. Second, at pitching, permutation analysis showed small but interpretable effects led by geography and financing: country_encoded 5.6%, capital_seeking_numeric 1.4%, Market_Opportunity 0.6%, with has_digital_health and has_website at 0.3% each and most remaining features ~ 0; the magnitudes are modest, consistent with rising uncertainty and sparse mid-funnel data. Third, due diligence elevates market realism, team execution, and defensible differentiation, aligning with committee emphasis on holistic defensibility; sector and stage context contribute but do not dominate. Performance follows the same arc: moderate discrimination at prescreen (AUC 0.717), comparable at pitching (AUC 0.692), and strongest at due diligence (AUC 0.836), with prospective pass-through showing that stringent final thresholds meaningfully constrain workload. Together, these results support a practical view from the literature: multimodal analysis is necessary to capture real decision signal from charts and tables, and governance is essential for trust and deployment (Guo et al. 2024 ; Yanglet, Cao, and Deng 2025). This work advances the field in three ways. Methodologically, it offers a reproducible path from interviews and questionnaires to an operational rubric and AI-derived embedding that respects leakage boundaries and test-set selection for honest reporting (Vollmer et al. 2020 ). Operationally, it turns thresholds into explicit policy levers rather than model afterthoughts, allowing partners to tune precision–recall to team capacity and market cadence while keeping a human-in-the-loop override. Governance-wise, it treats citations as first-class outputs, creating an audit trail that supports LP reporting and emerging compliance expectations in finance, thereby transforming a “black box” into a documented collaborator (Guo et al. 2024 ). Important limitations remain. Mid-funnel data sparsity constrains permutation-based interpretability for pitching and widens uncertainty around small effects; additional labeled cases will tighten estimates. Despite removing timing and future-dependent features, surrogate socio-geographic variables (for example, country encoding) can transmit historical bias; monitoring, fairness audits, and potential reweighting are warranted. Our prospective reporting used independent stage evaluations for transparency; fully progressive operational evaluation should be adopted to align pass-through dynamics with policy. Finally, models reflect NuFund’s thesis; transfer to other contexts requires re-eliciting local criteria and retraining, even though the methodology itself is portable. Next steps follow directly from the literature and our results. A preregistered, prospective validation on new cohorts should quantify effects on time-to-decision, review load, and hit rate while measuring reviewer agreement, calibration, and bias reduction under human–AI collaboration (Guo et al. 2024 ). An active-learning loop can route high-uncertainty deals for partner adjudication and fold expert feedback back into training to reduce drift over time. Extending beyond decks to structured “data room” analysis (legal, financial, technical) under stricter provenance checks will broaden diligence coverage without sacrificing explainability. Finally, as regulators formalize expectations around explainability and model risk, the disciplined practices used here—leakage controls, test-set reporting, and evidence-linked outputs—offer a template for trustworthy AI deployment in private markets (Guo et al. 2024 ; Vollmer et al. 2020 ). Methods We developed and evaluated a tailored AI framework to support venture capital deal screening in a progressive investment funnel. The framework integrates (1) investor-elicited indicators, (2) a domain-adapted AI analysis layer that produces a structured, interpretable embedding for each deal, and (3) supervised machine-learning (ML) models for prescreen, pitching, and due diligence stages. The instrument is organization-specific; the methodology is fully described below. Survey and Interview Protocol We used a two-phase elicitation process to derive the decision rubric and training signals. Semi-structured interviews: Experienced investors and reviewers participated in guided interviews covering team, market, product, traction, competition, financials, deal terms, and risk. Sessions were recorded in analysis notes and synthesized into thematic indicators and red-flag taxonomies. Structured questionnaire: A standardized instrument captured item-level preferences and threshold heuristics. Sections included: (A) Team/Execution, (B) Market/Go-to-Market, (C) Product/Defensibility, (D) Traction/Validation, (E) Financials/Forecasts, (F) Deal Terms/Use of Funds, (G) Risk Flags, and (H) Open-ended investment considerations. Closed-ended items used Likert scales (e.g., 1–5 agreement/importance), with optional free-text rationales. The instrument also included forced-choice tradeoffs (e.g., team strength vs. current traction) to surface threshold behavior. Synthesis to rubric: Interview themes and questionnaire distributions were mapped to a canonical rubric with score bands and flag definitions. We prioritized consistency and interpretability over granularity to support downstream model calibration and human review. Indicator Elicitation and Design Inputs We conducted semi-structured interviews and design workshops with investment committee members and active reviewers to capture decision heuristics and information needs. From these sessions, we derived a rubric covering team, market, business model, competitive advantage, financials, validation, scalability, and general AI assessment, alongside explicit risk flags and contextual metadata. The rubric served as the canonical schema for all downstream AI outputs and ML features. AI Analysis Layer We use Google’s Gemini 2.5 Pro as the analysis model. The system performs multimodal reading of pitch decks (PDF visuals and extracted text) and fuses application context (industry descriptors, founder LinkedIn profiles, funding ask) to generate structured, rubric-aligned outputs. Model and analysis procedure: Multimodal ingestion of PDFs plus contextual fields from the application record. Multi-pass critique loops to reduce single-pass bias and explicitly separate “missing information” from “fundamental flaws.” Outputs include dimension-level sub-scores, evidence-backed rationales, and normalized risk flags with severity levels (critical/moderate/minor). Evidence retrieval and validation integrated in analysis: The system generates targeted queries from deck content and context, then calls the Google Custom Search API to retrieve titles/snippets/URLs for external evidence on market size, growth rates, competitors, and regulatory context. Results are rate-limited, deduplicated (by URL/domain), and prioritized by relevance and recency. Non-resolving or redirected links are discarded. Validation logic favors authoritative sources (industry reports, major research firms, reputable news, journals). Discrepancies between pitch claims and external evidence are flagged; where evidence is insufficient, claims are labeled “could not be independently verified.” Structured embedding construction: The AI’s rubric-aligned sub-scores, flag taxonomy, and validated context features are normalized to produce an interpretable AI-derived embedding per deal. This embedding is designed for downstream ML: stable feature names, bounded scales, and explicit leakage controls (no future-dependent features, no operational timing fields). In addition, we generate structured HTML executive summaries with inline evidence hyperlinks and convert them to print-ready PDFs. Embedding and Dataset Construction AI outputs are consolidated into a tabular dataset together with selected contextual variables. Qualitative assessments are converted to numeric features and categorical variables are harmonized. Industry is represented by binary flags (AI, fintech, biotech, deep tech, digital health, software) and an industry-complexity score. Referral signals are parsed from free-text to create binary and coarse-type indicators. Additional features include capital sought (numeric), website and LinkedIn presence, and ten rubric round scores mapped to named dimensions (Team Expertise, Market Opportunity, Business Model, Competitive Advantage, Financial Projections, Red Flags, Investment Readiness, Validation, Scalability, General AI Assessment). To mitigate bias and information leakage, we exclude features that encode future or downstream information or operational timing (e.g., days since submission, days to closing, realized commitments and ratios, final aggregate scores, aggregate red-flag counts, urgency indicators), and we constrain high-cardinality descriptors. Model Development: Progressive Funnel We train independent classifiers for each stage of the investment funnel—prescreen, pitching, and due diligence—using the standardized table as input features. Candidate models include Logistic Regression, Random Forest, Gradient Boosting, and SVC, trained with a fixed random seed for reproducibility (random_state = 42). Model selection per stage is based on test-set AUC. The “best individual model” is chosen using held-out performance. After selection, a production copy of the best model is trained on the full dataset for that stage (to maximize prospective performance) and saved for deployment. Thresholding and Decision Logic To align with screening goals, we apply stage-specific probability thresholds calibrated for practical selectivity: prescreen 0.80, pitching 0.60, and due diligence 0.50. We additionally report a “high-confidence” designation when a model-specific, specificity-optimized threshold is exceeded. While operational use follows a progressive workflow, we report both progressive and non-progressive pass rates for transparency. Bias and Leakage Controls We exclude features known to encode future information or operational bias (e.g., timing variables relative to submission or closing, realized commitments or ratios, aggregate scoring fields). High-cardinality descriptors are reduced or encoded conservatively; industry is represented via binary flags and an industry-complexity score. Referral signals are parsed using robust keywording to avoid brittle categorical expansions. Evaluation We report AUC as the primary metric on a held-out test set and include sensitivity, specificity, precision, recall, F1, and confusion matrices for the selected model at each stage. For interpretability, we employ model-specific feature importance methods: Random Forest : Out-of-bag (OOB) feature importance, which provides unbiased estimates of feature importance by measuring the decrease in prediction accuracy when features are permuted in out-of-bag samples Support Vector Classification (SVC) : Permutation importance calculated on the test set with 10 repeats to ensure robust estimates of feature importance Other models : Built-in feature importance methods where available (e.g., coefficient magnitudes for linear models, built-in importance for gradient boosting) Feature importance visualizations are generated as compact, publication-ready panels showing the top 10 most important features for each model. Reproducibility and Constraints We fix random seeds during experimentation and separate model selection (test set) from production training (full dataset) to avoid optimistic reporting. All implementations are in Python (3.12), using pandas/numpy/scikit-learn for data/ML, matplotlib/seaborn for figures, and PyMuPDF for PDF processing; Google’s Gemini SDK and Google APIs are used for AI and search. The methodology—expert elicitation → AI embedding → supervised modeling—is general and transferable to other organizations with appropriate retuning. Adoption elsewhere typically requires re-eliciting investor indicators and validating outputs against local decision criteria. Declarations Ethics approval and consent to participate We conducted semi-structured interviews with NuFund investment partners and screeners to elicit rubric criteria and workflow constraints. Interviews were approved internally by NuFund and did not require IRB/REC review because participants were engaged in their professional capacity and no sensitive personal data were collected. Prior to each interview, participants received an information sheet and provided informed consent (verbal). Participation was voluntary and uncompensated. Transcripts were de-identified before analysis; any quotations presented in the manuscript are anonymized. Acknowledgments The authors used OpenAI’s ChatGPT 5 and Google’s Gemini 2.5 Pro for literature search, language editing, and clarity improvements, and GitHub Copilot and Anthropic’s Claude for code improvement. All ideas, analyses, and results presented in this manuscript are the original work of the authors. Competing interests Serhat Pala is President of NuFund Venture Group. The authors used de-identified deal data provided by NuFund for this study. The analysis and interpretation were conducted independently, and NuFund had no input into the analysis, results, or preparation of this manuscript. The framework described herein is presented for research purposes and does not disclose or release the operational tools currently in use by NuFund. All other authors declare no competing interests. References Askari M (2025) Reliable but supervised: evaluating a generative AI-rubric model for consistent and fair assessment in postgraduate education. Assess Evaluation High Educ :1–18 Guo E, Gupta M, Deng J, Park YJ, Paget M, Naugler C (2024) Automated Paper Screening for Clinical Reviews Using Large Language Models: Data Analysis Study. J Med Internet Res 26:e48996. 10.2196/48996 Vollmer S, Mateen BA, Bohner G, Kiraly FJ, Ghani R, Jonsson P, Cumbers S et al (2020) Machine learning and artificial intelligence research for patient benefit: 20 critical questions on transparency, replicability, ethics, and effectiveness. BMJ 368:l6927. 10.1136/bmj.l6927 Yanglet X-Y, Liu Y, Cao, Li, Deng (2025) Multimodal financial foundation models (mffms): Progress, prospects, and challenges. arXiv preprint arXiv:2506.01973 Additional Declarations The authors declare potential competing interests as follows: Serhat Pala is President of NuFund Venture Group. The authors used de-identified deal data provided by NuFund for this study. The analysis and interpretation were conducted independently, and NuFund had no input into the analysis, results, or preparation of this manuscript. The framework described herein is presented for research purposes and does not disclose or release the operational tools currently in use by NuFund. All other authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7781525","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":528947122,"identity":"52419b1a-e57f-47ed-bb4d-643898980ff6","order_by":0,"name":"Amin Haghani","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/klEQVRIiWNgGAWjYDACCQYGZiDFww9hg0ACkEuMFskGUrUwGBwgVgv/7ObHrwtq7sgY3z788DYPw2F5+fYExgdv2/BYcueYmfWMY894zM6lGVsDtRhuOPOA2XAuHi0MNxLMjHnYDvOYnWEwkwZqSTCQSGCT5sWjRf5G+jdjnn+HeYx72L+BtcjPSGD/jU+LwY0c48e8bYd5DHh4ILYA7WVjxqfF8EZOGTNv32EeiTM8xZZzDNKBfnnYLDnnHG4tcjfSN3/m+XbYnr+HfeONNxXWwBBLPvjhTRke7zMwsEnAWEw8BiCKsQGveiBg/gBjMf4gpHYUjIJRMApGJAAA9cBN0jvYMFwAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-6052-8793","institution":"Rady School of Management UCSD","correspondingAuthor":true,"prefix":"","firstName":"Amin","middleName":"","lastName":"Haghani","suffix":""},{"id":528947123,"identity":"f5340e92-18b6-47f8-92fb-3d84a1c0764d","order_by":1,"name":"Haydar Nur","email":"","orcid":"","institution":"Nufund Venture Group","correspondingAuthor":false,"prefix":"","firstName":"Haydar","middleName":"","lastName":"Nur","suffix":""},{"id":528947124,"identity":"a04feedc-25a6-42c9-b30d-713c816b5bc1","order_by":2,"name":"Serhat Pala","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9ElEQVRIiWNgGAWjYHACNiC2QeIzH2Bg4CGoJSENmZ9AlJbDJGiRb+999uDnj/Py5u3tjz98qLGRN2fjMXvwhsFOTrcBuxaDM8fNDXsSbhvOOXPGTHLGsTTDnW085oZzGJKNzQ7g0CKRxibBk3A7QUIih42Zt+Ew44b7PWbSPAwHErfh0CI//xmb5J+Ec0At6Y8/8zb8t99wjAe/FoYbbGzSPAkHgFoSDKR5Gw4kEtRicCaNTVomLdlwBg/YL8nJG46xlUnOMcDtF/n2Y2ySb2zs5CXYwSFmZ7vhGPM2iTcVdnK4tOACBqQpHwWjYBSMglGACgBuQFXl/YqP+AAAAABJRU5ErkJggg==","orcid":"","institution":"Nufund Venture Group","correspondingAuthor":true,"prefix":"","firstName":"Serhat","middleName":"","lastName":"Pala","suffix":""}],"badges":[],"createdAt":"2025-10-04 18:39:50","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":true,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-7781525/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7781525/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":94200616,"identity":"55f9da42-f991-4a56-ac6f-ae76003b2e1f","added_by":"auto","created_at":"2025-10-23 13:53:27","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":2202037,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscriptwithauthordetails.docx","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/b5ab3dd47f5dfcdd848822e6.docx"},{"id":94200619,"identity":"1da0ad53-424b-4620-996f-5c6ed2e99db7","added_by":"auto","created_at":"2025-10-23 13:53:28","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":342,"visible":true,"origin":"","legend":"","description":"","filename":"rs7781525.json","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/a53bede9f98085ffb4fc6be5.json"},{"id":94200674,"identity":"412c0029-2038-4daf-a618-921c06d796d6","added_by":"auto","created_at":"2025-10-23 13:53:34","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":50236,"visible":true,"origin":"","legend":"","description":"","filename":"rs77815250enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/57f3d70579e8c482d0aa628b.xml"},{"id":94200658,"identity":"094f58c1-19d7-4cce-a958-1ebb4dc6f7b1","added_by":"auto","created_at":"2025-10-23 13:53:32","extension":"emf","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":4677152,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.emf","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/3ffe522646f6776149c8c5ef.emf"},{"id":94200680,"identity":"53fa8db6-eee6-4688-89ec-9a0b2c2ead50","added_by":"auto","created_at":"2025-10-23 13:53:34","extension":"emf","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1642596,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.emf","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/4d4dbcca20b329c519999d02.emf"},{"id":94200673,"identity":"c60ca916-d4a8-46fc-8593-e7eb3035b6fc","added_by":"auto","created_at":"2025-10-23 13:53:34","extension":"emf","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1263596,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.emf","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/b86336eae78edaa23192df31.emf"},{"id":94200679,"identity":"6723a373-7fd1-4499-9b90-6056527d0d3c","added_by":"auto","created_at":"2025-10-23 13:53:34","extension":"emf","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1453736,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.emf","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/fada389c16d090c088fd092d.emf"},{"id":94200668,"identity":"8d3f7df4-ae3b-4b01-a340-041d7ea91d58","added_by":"auto","created_at":"2025-10-23 13:53:33","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":57954,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/e60d3ae0698c2aa5fea42cfa.png"},{"id":94200682,"identity":"e637f3b6-8643-4064-b371-6a9111659345","added_by":"auto","created_at":"2025-10-23 13:53:35","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":18454,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/a47025d39c1bf19b640f0a2d.png"},{"id":94200671,"identity":"d968ad5d-de4e-4103-a4af-e52a6ea6f0b9","added_by":"auto","created_at":"2025-10-23 13:53:34","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":16226,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/c8f3ac7715a9ed73a16d9e43.png"},{"id":94200686,"identity":"aa815b18-d26b-4270-a6a1-85189d2a23d3","added_by":"auto","created_at":"2025-10-23 13:53:35","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":22921,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/77c43c77ad4a3243f402297e.png"},{"id":94200662,"identity":"147eb4bf-a26e-42be-9272-1b166769d97d","added_by":"auto","created_at":"2025-10-23 13:53:33","extension":"xml","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":48697,"visible":true,"origin":"","legend":"","description":"","filename":"rs77815250structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/f2cb00690d4946361e8721ef.xml"},{"id":94200676,"identity":"d2d2f5c6-7577-4335-9cb6-da94513efff3","added_by":"auto","created_at":"2025-10-23 13:53:34","extension":"html","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":55340,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/21a1357410fc8f78e69519dd.html"},{"id":94200664,"identity":"eb720131-aeb1-4ea3-9386-d2ba2ab3629d","added_by":"auto","created_at":"2025-10-23 13:53:33","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":6190010,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 1.\u003c/strong\u003e \u003cstrong\u003eNuFund AI Deal Analysis Framework Workflow.\u003c/strong\u003e The comprehensive workflow diagram illustrates the integrated AI-powered venture capital deal analysis system. The framework begins with the Human Survey \u0026amp; AI Training component at the top, which continuously refines the system through feedback loops. Input data flows from multiple sources including pitch decks, application forms, and external market research into the AI Analysis Core, which processes information using Google's Gemini 2.5 Pro model. The AI generates structured embeddings and rubric scores that feed into the Embedding Hub, where 10-dimensional feature vectors are created for each deal. These embeddings support three main applications: the ML Pipeline for progressive funnel modeling (prescreen, pitching, due diligence), Executive Summaries with evidence-backed analysis, and the Due Diligence Toolkit for comprehensive deal evaluation. The QA \u0026amp; Validation component ensures data quality and consistency. Dashed feedback arrows from all applications back to the Human Survey component complete the continuous learning loop, enabling iterative improvement of the AI models based on investment outcomes and human expert feedback.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/e8a4f9837e9c11825874fb23.png"},{"id":94200660,"identity":"3610f70a-3219-4cf3-aaaf-31fc8728172e","added_by":"auto","created_at":"2025-10-23 13:53:33","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1003037,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 2.\u003c/strong\u003e \u003cstrong\u003eFunnel and Deal Characteristics.\u003c/strong\u003e \u003cstrong\u003eA)\u003c/strong\u003e Deal funnel trajectory showing the progressive flow of deals through the investment pipeline, from early stage (281 deals, 57.7%) through pre-screening (22 deals, 4.5%), pitching (2 deals, 0.4%), due diligence (6 deals, 1.2%), to final funding (77 deals, 15.8%), with rejected deals (80 deals, 16.4%). The funnel demonstrates the selective nature of venture capital investment, with significant attrition at each stage. \u003cstrong\u003eB)\u003c/strong\u003e Industry distribution analysis with two panels: left panel shows industry concentration across the deal portfolio (healthcare/biotech 44.0%, technology/AI 30.0%, other sectors 21.4%, finance and manufacturing/industrial each 2.3%), and right panel shows success rates by industry (% of deals reaching due diligence stage). \u003cstrong\u003eC)\u003c/strong\u003e Essential deal characteristics including round size distribution (capital seeking in millions of dollars, median 1.28M), investment stage distribution (1.28M), investment stage distribution (0.75M), providing insights into the financial profile and maturity of companies in the portfolio.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/6aa5a11077f5255ed4d759e4.png"},{"id":94200610,"identity":"ec0465c1-6638-4d9c-8d0d-e49955350a8a","added_by":"auto","created_at":"2025-10-23 13:53:26","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":813568,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 3.\u003c/strong\u003e \u003cstrong\u003eAI-Generated Embedding Analysis.\u003c/strong\u003e \u003cstrong\u003eA)\u003c/strong\u003e Box plots showing the distribution of AI embedding scores by funnel stage for five key assessment dimensions (Team Expertise, Market Opportunity, Business Model, Competitive Advantage, Financial Projections), with dynamic y-axis scaling to highlight score variations across stages. \u003cstrong\u003eB)\u003c/strong\u003e Multi-panel histograms displaying the frequency distribution of normalized AI embedding scores (0-10 scale) colored by funnel stage, showing how different stages cluster around specific score ranges for each assessment dimension. \u003cstrong\u003eC)\u003c/strong\u003e Correlation heatmap of the five key AI embedding features, revealing the relationships between different assessment dimensions and identifying which scores tend to co-occur in deal evaluations. \u003cstrong\u003eD)\u003c/strong\u003e Radar chart showing mean AI embedding scores by funnel stage, with scores ranging from 7-10, illustrating how different stages exhibit distinct patterns across the five assessment dimensions and highlighting the progression of deal quality through the investment funnel.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/a7297ca8a836379bf7e6648f.png"},{"id":94200666,"identity":"bb022313-564a-433b-8e8c-5192b7243334","added_by":"auto","created_at":"2025-10-23 13:53:33","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":819748,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFigure 4.\u003c/strong\u003e \u003cstrong\u003eMachine Learning Model Performance.\u003c/strong\u003e \u003cstrong\u003eA)\u003c/strong\u003e AUC comparison across five machine learning models (Random Forest, Support Vector Classifier, Logistic Regression, Gradient Boosting, and Ensemble) for each funnel stage (prescreen, pitching, due diligence), highlighting the best performing model for each stage. \u003cstrong\u003eB)\u003c/strong\u003e ROC curves for each funnel stage, showing the true positive rate versus false positive rate trade-off for the best performing model in each stage, with AUC scores displayed in the legend. \u003cstrong\u003eC)\u003c/strong\u003e Feature importance analysis for the best performing model in each stage, using permutation importance for Support Vector Classifier and built-in feature importance for Random Forest and Gradient Boosting models, displaying the top 10 most influential features for deal classification at each funnel stage. \u003cstrong\u003eD)\u003c/strong\u003e Sensitivity and specificity analysis for each funnel stage, showing the trade-off between true positive rate (sensitivity) and true negative rate (specificity) across different classification thresholds for the best performing model in each stage.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/b2fcb7c752036bd5a6a6a5a3.png"},{"id":94201632,"identity":"ed838fb5-69ed-457a-b1bf-fbf7cb4bbf67","added_by":"auto","created_at":"2025-10-23 14:01:23","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":6478706,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7781525/v1/056d77e1-7abf-4933-8891-1f7baefcae92.pdf"}],"financialInterests":"The authors declare potential competing interests as follows: Serhat Pala is President of NuFund Venture Group. The authors used de-identified deal data provided by NuFund for this study. The analysis and interpretation were conducted independently, and NuFund had no input into the analysis, results, or preparation of this manuscript. The framework described herein is presented for research purposes and does not disclose or release the operational tools currently in use by NuFund. All other authors declare no competing interests.","formattedTitle":"From Decks to Decisions: An Auditable AI Framework for Venture Capital","fulltext":[{"header":"Introduction","content":"\u003cp\u003eVenture capital (VC) faces a paradox of scale and scrutiny: record volumes of inbound opportunities, increasingly complex materials (pitch decks, KPI tables, product roadmaps), and heightened expectations for rigor and fairness in selection. Manual, network-driven screening strains under this load and can be inconsistent. At the same time, recent advances in large language models (LLMs)\u0026mdash;especially multimodal models capable of reading text, tables, and figures\u0026mdash;have shown promise in document triage and synthesis when paired with disciplined evidence and governance practices (Guo et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Yanglet, Cao, and Deng 2025).\u003c/p\u003e\u003cp\u003eMost prior approaches to VC screening are either generic (optimizing against public proxies rather than a fund\u0026rsquo;s thesis) or opaque (limiting adoption in regulated contexts). Text-centric pipelines also miss the visual and structural content that conveys signal in pitch decks (e.g., traction charts, unit-economics tables, product diagrams). Methodological pitfalls\u0026mdash;data leakage, look-ahead bias\u0026mdash;further inflate performance estimates, and black-box systems inhibit trust where provenance and explainability are required (Vollmer et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). By contrast, rubric-aligned outputs\u0026mdash;where the AI scores against predefined criteria with rationales\u0026mdash;improve interpretability and reliability for downstream decisions and assessments (Askari \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eWe introduce an applied, auditable AI framework that addresses these gaps by aligning analysis and modeling to a fund-specific investment rubric. First, we elicit investor criteria through interviews and a structured questionnaire to encode a rubric that reflects the committee\u0026rsquo;s actual decision factors. Second, a multimodal model (Google Gemini 2.5 Pro) reads pitch decks and contextual materials and produces a structured embedding\u0026mdash;criterion-level scores, rationales, and direct evidence links\u0026mdash;enforcing a stable schema and reducing free-text variance (Guo et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Yanglet, Cao, and Deng 2025). Third, we transform this embedding into features for a three-stage machine-learning funnel (prescreen, pitching, due diligence) with strict leakage controls and test-set model selection to ensure honest evaluation (Vollmer et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Finally, we generate investor-facing HTML summaries with inline citations and print-ready PDFs under a governance regime emphasizing provenance, auditability, and reproducibility.\u003c/p\u003e\u003cp\u003eContributions are threefold. (i) A multimodal, rubric-aligned analysis layer that converts heterogeneous pitch materials into verifiable, structured signals. (ii) A leakage-aware ML pipeline tailored to VC\u0026rsquo;s progressive funnel, with stage-specific thresholds calibrated as operational policy levers. (iii) A governed path from raw decks to decision support that foregrounds explainability and evidence, suitable for peer review and operational deployment (Vollmer et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Askari \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2025\u003c/span\u003e; Guo et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Yanglet, Cao, and Deng 2025). We position this as an exemplar implementation at NuFund; the methodology is transferable, though adoption elsewhere requires re-eliciting investor indicators and validating outputs against local decision criteria.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eHuman Survey \u0026amp; AI Training Loop\u003c/p\u003e\u003cp\u003eWe implemented a continuous human feedback mechanism that enables ongoing refinement of the AI system based on investor evaluations. This component collects structured feedback from NuFund investors on AI-generated predictions and reports, allowing for rubric refinement, threshold calibration, and model improvement. The feedback loop ensures the AI system remains aligned with evolving investor preferences and market conditions, creating a dynamic learning system that improves over time.\u003c/p\u003e\u003cp\u003eInvestor Survey and Rubric Elicitation\u003c/p\u003e\u003cp\u003eWe conducted semi-structured interviews and a structured questionnaire to elicit decision criteria and operational cutoffs from NuFund investors. The final rubric distilled recurring indicators across market, team, product, traction, and financing terms into bounded, interpretable sub-scores and risk flags, which the AI layer uses to generate a structured, evidence-linked embedding for each deal (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). This process standardized language across reviewers and enabled consistent application of stage-specific thresholds downstream. The initial survey insights now feed into the continuous Human Survey \u0026amp; AI Training loop for ongoing refinement.\u003c/p\u003e\u003cp\u003eKey findings from interviews (synthesized):\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eFounders and team\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eStrong founder\u0026ndash;market fit and complementary founding teams were universally emphasized; solo founders at POC are higher risk.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIntegrity and commitment checks (fraud/bankruptcy/legal history; time dedication) are mandatory screens; inflated salaries at early stage are a red flag.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eMarket and competition\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eValidated market research and credible SAM/TAM assumptions are required; over-broad or inflated markets are a red flag.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eStrategic small markets can be acceptable if there is clear strategic importance; otherwise small/poorly defined markets are disfavored.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eProduct and roadmap\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eScalability and feasibility must be evidenced; over-customization and unrealistic roadmaps are negatives.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eClear differentiation and awareness of direct/indirect competitors are expected.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eFinancials and terms\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eRealistic projections and clean cap tables are baseline; misleading metrics (e.g., bookings vs ARR) or heavy debt are red flags.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eCapital efficiency and a credible cash-flow/raise plan are positives; dependency on large unfunded raises is risky.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eLegal/regulatory and presentation\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eEarly regulatory strategy (e.g., FDA) and IP posture should be addressed in regulated sectors.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eDeck quality matters: concise, professional, error-free, with defensible numbers; cluttered/error-filled decks signal risk.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eDistinct emphases by investor\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eRisk focusing (avoid taking high market and high product risk simultaneously); commercialization plans in med-device; diversity as a team strength; skepticism of AI over-claims without technical depth; acquisition potential in life sciences.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThese insights directly informed the 24-criterion rubric (\u0026minus;\u0026thinsp;10 to +\u0026thinsp;10 per item) spanning founders, market, product, financials, regulatory, and presentation, yielding a total score range of \u0026minus;\u0026thinsp;240 to +\u0026thinsp;240 with qualitative interpretation bands for screening (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eDatabase and Funnel Characteristics\u003c/p\u003e\u003cp\u003eOur Dealum cohort comprised 487 unique companies. Most submissions were early-stage: 57.7% at intake (Applied/Referral), while 15.8% were ultimately Funded and 16.4% Rejected/Archived. Industry composition (known for 80.7% of deals) was led by Healthcare/Biotech (44.0%) and Technology/AI (30.0%). Companies typically sought modest raises (median 1.28M USD; mean 2.13M USD; range 7.25K USD to 65.0M USD) and were concentrated at pre-seed/seed (76.0% of those reporting stage). Prior funding was sparse (median 0.75M USD; mean 2.34M USD). Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e summarizes the funnel, top industries, and capital/stage/funding distributions.\u003c/p\u003e\u003cp\u003eThe progressive workflow comprises three stages\u0026mdash;prescreen, pitching, and due diligence\u0026mdash;each calibrated to shape workload and selectivity. On an independent prospective set of 281 new deals, high-confidence thresholds yielded: prescreen 59.1% pass, 10.2% of prescreen passes advancing to pitching (6.0% of total), and 0.0% advancing to due diligence (reflecting stringent final-stage selectivity). For operational practice, we also used relaxed decision thresholds (0.8, 0.6, 0.5 for prescreen, pitching, due diligence) to flag candidates for review while reserving 0.823/0.685/0.704 as high-confidence cutoffs.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eAI-Generated Embedding Analysis\u003c/p\u003e\u003cp\u003eThe AI system generates structured embeddings for each deal based on the 24-criterion rubric, creating 10-dimensional feature vectors that capture key assessment dimensions: Team Expertise, Market Opportunity, Business Model, Competitive Advantage, and Financial Projections. These embeddings are normalized to a 0\u0026ndash;10 scale and provide quantitative representations of deal quality across multiple dimensions. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows the distribution of these AI-generated scores across funnel stages, revealing distinct patterns in how different stages cluster around specific score ranges. The correlation analysis demonstrates relationships between assessment dimensions, while the radar chart illustrates how mean scores vary systematically across funnel stages, highlighting the progression of deal quality through the investment process.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eProgressive ML Performance\u003c/p\u003e\u003cp\u003eModels were selected by test-set AUC per stage. The best individual models were GradientBoosting (prescreen, AUC 0.717), SVC (pitching, AUC 0.692), and SVC (due diligence, AUC 0.836). Thresholds used operational cutoffs aligned to workload and selectivity (prescreen 0.823; pitching 0.685; due diligence 0.704 for high-confidence reporting). Stage-level ROC, PR, confusion matrix, and sensitivity/specificity panels are provided in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e (A\u0026ndash;D) and supplemental panels.\u003c/p\u003e\u003cp\u003eFeature importance and feature set\u003c/p\u003e\u003cp\u003eAll stages use an identical 25-feature set (15 structured variables such as industry flags, stage and geography encodings, capital seeking, web presence; and 10 AI-derived rubric scores spanning team, market, product, finance, and scalability).\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003ePrescreen (GradientBoosting): The most influential features were country_encoded (24.2%), Team_Expertise (20.1%), capital_seeking_numeric (8.4%), Financial_Projections (8.3%), and company_stage_encoded (7.9%), followed by Competitive_Advantage and Market_Opportunity (~\u0026thinsp;6\u0026ndash;7%). This mirrors interview priorities that emphasized strong teams, realistic funding asks/projections, and stage fit.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003ePitching (SVC): Same feature set; the figure shows country_encoded (5.6%) as the top feature, followed by capital_seeking_numeric (1.4%), Market_Opportunity (0.6%), has_digital_health (0.3%), and has_website (0.3%); most other features show zero importance. This aligns with prescreen patterns where geographic and financial factors dominate early-stage decisions, though with much lower overall importance than the prescreen stage.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eDue diligence (SVC, permutation importance): Market_Opportunity was the top signal (~\u0026thinsp;3.5%), followed by Team_Expertise (~\u0026thinsp;3.2%) and Competitive_Advantage (~\u0026thinsp;2.8%); sector and stage signals were smaller (has_digital_health\u0026thinsp;~\u0026thinsp;2.5%, funding_stage_encoded\u0026thinsp;~\u0026thinsp;2.2%). This aligns with interviews that final decisions emphasize market realism, team execution capacity, and defensible differentiation, with sector/stage context contributing but not dominating.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eOur findings demonstrate that a governed, rubric-aligned, multimodal AI layer can translate heterogeneous pitch materials into auditably structured signals that are both usable by downstream ML and verifiable by investors. In contrast to much of the startup-prediction literature that relies on generic public features and text-only embeddings\u0026mdash;and is often weakened by look-ahead bias or leakage\u0026mdash;our approach pairs an investor-elicited rubric with a multimodal model to capture visual and tabular signal from decks, then enforces verifiable provenance through inline citations (Guo et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Vollmer et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Yanglet, Cao, and Deng 2025). This design directly addresses two gaps highlighted in prior work and practice: the absence of fund-specific criteria in modeling pipelines and the opacity of \u0026ldquo;black-box\u0026rdquo; systems that hinder adoption in regulated contexts. By forcing the AI to produce rubric-aligned scores, rationales, and evidence links, we shift from free-text outputs to an interpretable schema that improves both machine learning stability and human review, consistent with evidence that rubric-aligned assessments enhance transparency and reliability (Askari \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eEmpirically, three stage-wise patterns connect investor judgment with model behavior. First, at prescreen, geographic context and team capacity dominate alongside pragmatic financing signals; this mirrors interview themes around founder-market fit, execution readiness, and credible capital plans. Second, at pitching, permutation analysis showed small but interpretable effects led by geography and financing: country_encoded 5.6%, capital_seeking_numeric 1.4%, Market_Opportunity 0.6%, with has_digital_health and has_website at 0.3% each and most remaining features\u0026thinsp;~\u0026thinsp;0; the magnitudes are modest, consistent with rising uncertainty and sparse mid-funnel data. Third, due diligence elevates market realism, team execution, and defensible differentiation, aligning with committee emphasis on holistic defensibility; sector and stage context contribute but do not dominate. Performance follows the same arc: moderate discrimination at prescreen (AUC 0.717), comparable at pitching (AUC 0.692), and strongest at due diligence (AUC 0.836), with prospective pass-through showing that stringent final thresholds meaningfully constrain workload. Together, these results support a practical view from the literature: multimodal analysis is necessary to capture real decision signal from charts and tables, and governance is essential for trust and deployment (Guo et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Yanglet, Cao, and Deng 2025).\u003c/p\u003e\u003cp\u003eThis work advances the field in three ways. Methodologically, it offers a reproducible path from interviews and questionnaires to an operational rubric and AI-derived embedding that respects leakage boundaries and test-set selection for honest reporting (Vollmer et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Operationally, it turns thresholds into explicit policy levers rather than model afterthoughts, allowing partners to tune precision\u0026ndash;recall to team capacity and market cadence while keeping a human-in-the-loop override. Governance-wise, it treats citations as first-class outputs, creating an audit trail that supports LP reporting and emerging compliance expectations in finance, thereby transforming a \u0026ldquo;black box\u0026rdquo; into a documented collaborator (Guo et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eImportant limitations remain. Mid-funnel data sparsity constrains permutation-based interpretability for pitching and widens uncertainty around small effects; additional labeled cases will tighten estimates. Despite removing timing and future-dependent features, surrogate socio-geographic variables (for example, country encoding) can transmit historical bias; monitoring, fairness audits, and potential reweighting are warranted. Our prospective reporting used independent stage evaluations for transparency; fully progressive operational evaluation should be adopted to align pass-through dynamics with policy. Finally, models reflect NuFund\u0026rsquo;s thesis; transfer to other contexts requires re-eliciting local criteria and retraining, even though the methodology itself is portable.\u003c/p\u003e\u003cp\u003eNext steps follow directly from the literature and our results. A preregistered, prospective validation on new cohorts should quantify effects on time-to-decision, review load, and hit rate while measuring reviewer agreement, calibration, and bias reduction under human\u0026ndash;AI collaboration (Guo et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e). An active-learning loop can route high-uncertainty deals for partner adjudication and fold expert feedback back into training to reduce drift over time. Extending beyond decks to structured \u0026ldquo;data room\u0026rdquo; analysis (legal, financial, technical) under stricter provenance checks will broaden diligence coverage without sacrificing explainability. Finally, as regulators formalize expectations around explainability and model risk, the disciplined practices used here\u0026mdash;leakage controls, test-set reporting, and evidence-linked outputs\u0026mdash;offer a template for trustworthy AI deployment in private markets (Guo et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Vollmer et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eWe developed and evaluated a tailored AI framework to support venture capital deal screening in a progressive investment funnel. The framework integrates (1) investor-elicited indicators, (2) a domain-adapted AI analysis layer that produces a structured, interpretable embedding for each deal, and (3) supervised machine-learning (ML) models for prescreen, pitching, and due diligence stages. The instrument is organization-specific; the methodology is fully described below.\u003c/p\u003e\u003cp\u003eSurvey and Interview Protocol\u003c/p\u003e\u003cp\u003eWe used a two-phase elicitation process to derive the decision rubric and training signals.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eSemi-structured interviews: Experienced investors and reviewers participated in guided interviews covering team, market, product, traction, competition, financials, deal terms, and risk. Sessions were recorded in analysis notes and synthesized into thematic indicators and red-flag taxonomies.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eStructured questionnaire: A standardized instrument captured item-level preferences and threshold heuristics. Sections included: (A) Team/Execution, (B) Market/Go-to-Market, (C) Product/Defensibility, (D) Traction/Validation, (E) Financials/Forecasts, (F) Deal Terms/Use of Funds, (G) Risk Flags, and (H) Open-ended investment considerations. Closed-ended items used Likert scales (e.g., 1\u0026ndash;5 agreement/importance), with optional free-text rationales. The instrument also included forced-choice tradeoffs (e.g., team strength vs. current traction) to surface threshold behavior.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eSynthesis to rubric: Interview themes and questionnaire distributions were mapped to a canonical rubric with score bands and flag definitions. We prioritized consistency and interpretability over granularity to support downstream model calibration and human review.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eIndicator Elicitation and Design Inputs\u003c/p\u003e\u003cp\u003eWe conducted semi-structured interviews and design workshops with investment committee members and active reviewers to capture decision heuristics and information needs. From these sessions, we derived a rubric covering team, market, business model, competitive advantage, financials, validation, scalability, and general AI assessment, alongside explicit risk flags and contextual metadata. The rubric served as the canonical schema for all downstream AI outputs and ML features.\u003c/p\u003e\u003cp\u003eAI Analysis Layer\u003c/p\u003e\u003cp\u003eWe use Google\u0026rsquo;s Gemini 2.5 Pro as the analysis model. The system performs multimodal reading of pitch decks (PDF visuals and extracted text) and fuses application context (industry descriptors, founder LinkedIn profiles, funding ask) to generate structured, rubric-aligned outputs.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eModel and analysis procedure:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eMultimodal ingestion of PDFs plus contextual fields from the application record.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eMulti-pass critique loops to reduce single-pass bias and explicitly separate \u0026ldquo;missing information\u0026rdquo; from \u0026ldquo;fundamental flaws.\u0026rdquo;\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eOutputs include dimension-level sub-scores, evidence-backed rationales, and normalized risk flags with severity levels (critical/moderate/minor).\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eEvidence retrieval and validation integrated in analysis:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eThe system generates targeted queries from deck content and context, then calls the Google Custom Search API to retrieve titles/snippets/URLs for external evidence on market size, growth rates, competitors, and regulatory context.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eResults are rate-limited, deduplicated (by URL/domain), and prioritized by relevance and recency. Non-resolving or redirected links are discarded.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eValidation logic favors authoritative sources (industry reports, major research firms, reputable news, journals). Discrepancies between pitch claims and external evidence are flagged; where evidence is insufficient, claims are labeled \u0026ldquo;could not be independently verified.\u0026rdquo;\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eStructured embedding construction:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eThe AI\u0026rsquo;s rubric-aligned sub-scores, flag taxonomy, and validated context features are normalized to produce an interpretable AI-derived embedding per deal.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eThis embedding is designed for downstream ML: stable feature names, bounded scales, and explicit leakage controls (no future-dependent features, no operational timing fields).\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIn addition, we generate structured HTML executive summaries with inline evidence hyperlinks and convert them to print-ready PDFs.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eEmbedding and Dataset Construction\u003c/p\u003e\u003cp\u003eAI outputs are consolidated into a tabular dataset together with selected contextual variables. Qualitative assessments are converted to numeric features and categorical variables are harmonized. Industry is represented by binary flags (AI, fintech, biotech, deep tech, digital health, software) and an industry-complexity score. Referral signals are parsed from free-text to create binary and coarse-type indicators. Additional features include capital sought (numeric), website and LinkedIn presence, and ten rubric round scores mapped to named dimensions (Team Expertise, Market Opportunity, Business Model, Competitive Advantage, Financial Projections, Red Flags, Investment Readiness, Validation, Scalability, General AI Assessment). To mitigate bias and information leakage, we exclude features that encode future or downstream information or operational timing (e.g., days since submission, days to closing, realized commitments and ratios, final aggregate scores, aggregate red-flag counts, urgency indicators), and we constrain high-cardinality descriptors.\u003c/p\u003e\u003cp\u003eModel Development: Progressive Funnel\u003c/p\u003e\u003cp\u003eWe train independent classifiers for each stage of the investment funnel\u0026mdash;prescreen, pitching, and due diligence\u0026mdash;using the standardized table as input features.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eCandidate models include Logistic Regression, Random Forest, Gradient Boosting, and SVC, trained with a fixed random seed for reproducibility (random_state\u0026thinsp;=\u0026thinsp;42).\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eModel selection per stage is based on test-set AUC. The \u0026ldquo;best individual model\u0026rdquo; is chosen using held-out performance.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eAfter selection, a production copy of the best model is trained on the full dataset for that stage (to maximize prospective performance) and saved for deployment.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThresholding and Decision Logic\u003c/p\u003e\u003cp\u003eTo align with screening goals, we apply stage-specific probability thresholds calibrated for practical selectivity: prescreen 0.80, pitching 0.60, and due diligence 0.50. We additionally report a \u0026ldquo;high-confidence\u0026rdquo; designation when a model-specific, specificity-optimized threshold is exceeded. While operational use follows a progressive workflow, we report both progressive and non-progressive pass rates for transparency.\u003c/p\u003e\u003cp\u003eBias and Leakage Controls\u003c/p\u003e\u003cp\u003eWe exclude features known to encode future information or operational bias (e.g., timing variables relative to submission or closing, realized commitments or ratios, aggregate scoring fields). High-cardinality descriptors are reduced or encoded conservatively; industry is represented via binary flags and an industry-complexity score. Referral signals are parsed using robust keywording to avoid brittle categorical expansions.\u003c/p\u003e\u003cp\u003eEvaluation\u003c/p\u003e\u003cp\u003eWe report AUC as the primary metric on a held-out test set and include sensitivity, specificity, precision, recall, F1, and confusion matrices for the selected model at each stage. For interpretability, we employ model-specific feature importance methods:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eRandom Forest\u003c/b\u003e: Out-of-bag (OOB) feature importance, which provides unbiased estimates of feature importance by measuring the decrease in prediction accuracy when features are permuted in out-of-bag samples\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eSupport Vector Classification (SVC)\u003c/b\u003e: Permutation importance calculated on the test set with 10 repeats to ensure robust estimates of feature importance\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eOther models\u003c/b\u003e: Built-in feature importance methods where available (e.g., coefficient magnitudes for linear models, built-in importance for gradient boosting)\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eFeature importance visualizations are generated as compact, publication-ready panels showing the top 10 most important features for each model.\u003c/p\u003e\u003cp\u003eReproducibility and Constraints\u003c/p\u003e\u003cp\u003eWe fix random seeds during experimentation and separate model selection (test set) from production training (full dataset) to avoid optimistic reporting. All implementations are in Python (3.12), using pandas/numpy/scikit-learn for data/ML, matplotlib/seaborn for figures, and PyMuPDF for PDF processing; Google\u0026rsquo;s Gemini SDK and Google APIs are used for AI and search. The methodology\u0026mdash;expert elicitation \u0026rarr; AI embedding \u0026rarr; supervised modeling\u0026mdash;is general and transferable to other organizations with appropriate retuning. Adoption elsewhere typically requires re-eliciting investor indicators and validating outputs against local decision criteria.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u0026nbsp;Ethics approval and consent to participate We conducted semi-structured interviews with NuFund investment partners and screeners to elicit rubric criteria and workflow constraints. Interviews were approved internally by NuFund and did not require IRB/REC review because participants were engaged in their professional capacity and no sensitive personal data were collected. Prior to each interview, participants received an information sheet and provided informed consent (verbal). Participation was voluntary and uncompensated. Transcripts were de-identified before analysis; any quotations presented in the manuscript are anonymized.\u003c/p\u003e\u003ch2\u003eAcknowledgments\u003c/h2\u003e\u003cp\u003eThe authors used OpenAI\u0026rsquo;s ChatGPT 5 and Google\u0026rsquo;s Gemini 2.5 Pro for literature search, language editing, and clarity improvements, and GitHub Copilot and Anthropic\u0026rsquo;s Claude for code improvement. All ideas, analyses, and results presented in this manuscript are the original work of the authors.\u003c/p\u003e\u003cp\u003eCompeting interests\u003c/p\u003e\u003cp\u003eSerhat Pala is President of NuFund Venture Group. The authors used de-identified deal data provided by NuFund for this study. The analysis and interpretation were conducted independently, and NuFund had no input into the analysis, results, or preparation of this manuscript. The framework described herein is presented for research purposes and does not disclose or release the operational tools currently in use by NuFund. All other authors declare no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAskari M (2025) Reliable but supervised: evaluating a generative AI-rubric model for consistent and fair assessment in postgraduate education. Assess Evaluation High Educ :1\u0026ndash;18\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGuo E, Gupta M, Deng J, Park YJ, Paget M, Naugler C (2024) Automated Paper Screening for Clinical Reviews Using Large Language Models: Data Analysis Study. J Med Internet Res 26:e48996. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2196/48996\u003c/span\u003e\u003cspan address=\"10.2196/48996\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVollmer S, Mateen BA, Bohner G, Kiraly FJ, Ghani R, Jonsson P, Cumbers S et al (2020) Machine learning and artificial intelligence research for patient benefit: 20 critical questions on transparency, replicability, ethics, and effectiveness. BMJ 368:l6927. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/bmj.l6927\u003c/span\u003e\u003cspan address=\"10.1136/bmj.l6927\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYanglet X-Y, Liu Y, Cao, Li, Deng (2025) Multimodal financial foundation models (mffms): Progress, prospects, and challenges. \u003cem\u003earXiv preprint arXiv:2506.01973\u003c/em\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"University of California San Diego","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Artificial Intelligence, LLM, Venture Capital, Machine Learning, Investment Screening, Explainable AI","lastPublishedDoi":"10.21203/rs.3.rs-7781525/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7781525/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eVenture capital faces a paradox of scale and scrutiny: rising volumes of inbound deals, increasingly multimodal pitch materials, and higher expectations for rigor and fairness. We present an auditable AI framework that converts pitch decks into a rubric-aligned, evidence-linked embedding using a multimodal model, then trains leakage-aware classifiers for a progressive funnel (prescreen, pitching, due diligence). On held-out tests, the best stage models achieved AUCs of 0.717 (prescreen), 0.692 (pitching), and 0.836 (due diligence). In a prospective cohort (n\u0026thinsp;=\u0026thinsp;281), high-confidence thresholds shaped workload as intended: 59.1% passed prescreen, 10.2% of those advanced to pitching (6.0% of total), and 0.0% advanced to due diligence, reflecting deliberate final-stage stringency. Feature profiles aligned with investor priorities: geography and team capacity at intake; market realism, team execution, and defensible differentiation at diligence; mid-funnel effects were smaller but interpretable. Unlike generic or opaque pipelines, our approach centers governance\u0026mdash;direct citations, strict leakage controls, test-set model selection\u0026mdash;turning thresholds into policy dials and AI into a documented collaborator. The methodology is transferable (re-elicitation required), and prospective validation with human\u0026ndash;AI collaboration studies and active-learning loops is a practical next step for trustworthy deployment in private markets.\u003c/p\u003e","manuscriptTitle":"From Decks to Decisions: An Auditable AI Framework for Venture Capital","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-23 13:53:14","doi":"10.21203/rs.3.rs-7781525/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7e8c702f-2621-4b35-9c6c-8613f6db13ea","owner":[],"postedDate":"October 23rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":56218373,"name":"Artificial Intelligence and Machine Learning"},{"id":56218374,"name":"Operations Research"},{"id":56218375,"name":"Finance"},{"id":56218376,"name":"Entrepreneurship"}],"tags":[],"updatedAt":"2025-10-23T13:53:14+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-23 13:53:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7781525","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7781525","identity":"rs-7781525","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.