Gen2: Building a Reviewer-Defensible Benchmark for Binding Hypothesis Triage in Cryptic Pocket Discovery | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Method Article Gen2: Building a Reviewer-Defensible Benchmark for Binding Hypothesis Triage in Cryptic Pocket Discovery Mr shakeel Hoosdally This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9336311/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Benchmarking in cryptic-pocket and allosteric discovery is often weakened by forcing heterogeneous case studies into pooled scoring despite ambiguous labels, unstable site assignment, missing row-level outputs, or mismatched evidential standards. Here, we present Gen2 as a governance-first benchmark framework for binding hypothesis triage in cryptic pocket discovery. Rather than treating benchmark assembly as a secondary administrative step, Gen2 treats it as part of the scientific method: each candidate slice is screened against frozen evidential rules, assigned a bounded role, and either admitted, parked, excluded, or retained as calibration or falsification material before pooled evaluation is considered. Applying this framework to the current panel produced a preserved no-active-slice-open checkpoint. Under these rules, HIF-2α remained policy-closed, TP53 Y220C remained calibration-only, CK2 was retained as falsification material, and KRAS G12D and PTP1B remained non-row-ready for different reasons. The principal result is therefore not pooled benchmark performance, but demonstration that Gen2 prevents invalid pooled claims by blocking premature scoring and preserving only reviewer-defensible evaluable units. This establishes a reproducible benchmark-construction layer for future multi-slice evaluation once row-ready systems and explicit row mappings are available. Computational Biology cryptic pocket allostery benchmark construction benchmark governance binding hypothesis triage slice admissibility pooled scoring falsification slice PTP1B TP53 Y220C KRAS G12D Figures Figure 1 Simple Summary Benchmark papers in cryptic-pocket discovery can become misleading when weak, ambiguous, or incomplete cases are forced into pooled scoring. In this study, Gen2 is presented as a benchmark framework that first decides which cases are actually fit for evaluation. Applied to the current panel, this process showed that several candidate slices should remain parked, calibration-only, or excluded rather than being scored prematurely. The main result is therefore not benchmark performance, but a controlled framework that prevents invalid pooled claims and defines what must be resolved before a full benchmark can be done. 1. Introduction Cryptic-pocket discovery remains methodologically difficult because the central object of interest is often not fully visible in the ligand-free structure. In many systems, a site that later accommodates a ligand is absent, occluded, or too weakly formed to be recognized confidently from a single apo conformation alone [ 1 ]. This has made cryptic pockets attractive in drug discovery, particularly for targets that appear poorly tractable in their ground-state structures, but it has also created a recurring evidential problem. Methods based on molecular dynamics, mixed-solvent sampling, or other ensemble-based strategies can reveal conformations that static analysis misses, yet they can also generate many transient cavities that are not credible small-molecule binding hypotheses [ 1 ]. The practical difficulty is therefore not only detecting pockets, but deciding which proposed ligand-pocket hypotheses are strong enough to count as benchmark evidence. That problem becomes sharper when benchmark claims are made. Recent work on cryptic-site datasets has highlighted a basic weakness in the field: performance can look artificially strong when methods are trained or evaluated mainly on ligand-bound states rather than on matched apo-holo settings, even though cryptic-site prediction is supposed to address precisely the gap between those states [ 2 ]. In other words, a benchmark can fail before any scoring begins if heterogeneous cases are pooled despite ambiguous labels, unstable site assignment, missing row-level outputs, or mismatched evidential standards. For cryptic-pocket and allosteric systems, this is not a minor bookkeeping issue. It directly affects whether pooled scoring reflects a real discrimination task or only a mixture of partially comparable cases. We recently introduced Gen2 as a biophysical triage framework for evaluating ligand-pocket hypotheses in cryptic-pocket discovery [ 3 ]. That earlier study focused on case-level adjudication: whether a proposed binding hypothesis remained credible after stricter biophysical scrutiny in individual systems. The present manuscript addresses a different question. Here, the focus is the benchmark-construction layer itself: how candidate slices are screened, how their roles are bounded, when they can be treated as evaluable, and when they must remain calibration-only, parked, policy-closed, or excluded from pooled scoring. Accordingly, this paper does not present a completed pooled benchmark and does not fold validation or assay layers into the benchmark core. Instead, it asks a narrower question: can Gen2 impose enough evidential discipline to distinguish reviewer-defensible benchmark material from non-evaluable material before pooled scoring is attempted? In this framework, benchmark curation, slice locking, table repair, and explicit marking of unevaluated rows are not background administration. They are part of the scientific result. 2. Materials and Methods 2.1. Benchmark claim and scope This study was designed as a benchmark-construction analysis rather than as a completed pooled benchmark or an experimental validation study. The purpose was to define how candidate slices enter the benchmark program, how their roles are assigned, and under what conditions they can be treated as evaluable material for later cross-slice comparison. Accordingly, the present Methods focus on benchmark curation, slice locking, row-status assignment, and pooled-metric eligibility. Validation layers such as blinded orthogonal reanalysis, biophysical assay follow-up, and assay-integration logic were kept outside the benchmark core by design. This separation was adopted because cryptic-site benchmarking is especially vulnerable to inflated performance when heterogeneous cases are pooled before their evidential status is made comparable [ 2 ]. 2.2. Prediction unit The basic unit of analysis was defined as one ligand-target-site hypothesis per row. Each row therefore represented a specific claim: that a named ligand, in a defined target system and structural context, supported or failed to support a defined binding-site hypothesis. This row-level design was chosen to prevent case-level narrative drift and to force explicit treatment of ligand identity, target state, proposed site, evidence class, and benchmark role before any result was counted. In practice, benchmark assembly could not proceed by describing a target in general terms alone. Instead, each proposed benchmark element had to be reducible to auditable row entries that could later be classified as evaluable, unevaluated, excluded, calibration-only, comparator-only, or otherwise restricted. 2.3. Ground-truth and label rules Row labels were assigned under bounded evidential rules. A positive row required direct or closely linked evidence supporting the stated ligand-target-site hypothesis. A negative row required explicit inactivity, no-effect behavior, or equivalent rejection evidence traceable to a primary source. Rows lacking sufficient support were not forced into positive or negative classes. Instead, they were retained as provisional, unresolved, or unevaluated until their evidential status could be clarified. This rule was intended to prevent false confidence from partial literature support, incomplete structural context, or weak comparator logic. Untested ligands, ambiguous site assignments, unresolved identities, and rows lacking a credible comparator were therefore retained inside the benchmark infrastructure only as non-pooled material. They could inform later slice development, but they were not allowed to contribute to pooled benchmark claims [ 8 – 12 ]. 2.4. Candidate-system screening and slice admission criteria Candidate systems were screened using a structural-first benchmark logic. A slice was considered suitable only if it presented a clear pocket-level adjudication question, at least one credible anchor positive, at least one fair comparator, adequate public structural support, and a task definition that could be handled under the existing decision framework without rewriting the charter. Systems were not promoted simply because they were biologically important or well represented in the literature. They were promoted only when the available evidence could be translated into a bounded benchmark problem with explicit row-level registration. Public structural and literature resources used during screening included the RCSB Protein Data Bank, PubMed, KLIFS, and ASD2023 [ 4 – 7 ]. This screening step also distinguished between benchmark roles. Some systems were suitable as minimal benchmark slices, some only as limited calibration or comparator material, and some only as falsification or failure-mode cases. Systems that lacked clean comparator logic, contained unresolved structural ambiguity, or could not be mapped to row-level adjudication without interpretive inflation were not admitted as pooled benchmark material. 2.5. Benchmark panel ontology Before any slice was interpreted, each candidate system was assigned a bounded benchmark role. The allowed roles were: included benchmark slice, limited calibration slice, limited comparator slice, controlled mature-slice buildout, falsification or failure-mode slice, and excluded slice. This ontology was introduced to prevent heterogeneous cases from being pooled under a single benchmark label when their evidential status was not actually comparable. In practical terms, it allowed the benchmark to distinguish between slices that could contribute to later scoring, slices that could only provide calibration or restricted comparator value, and slices that were useful only for documenting failure or rejection behavior. The purpose of this role system was not to make the panel appear larger, but to prevent weakly supported or mismatched cases from being promoted into benchmark evidence simply because they were available in the literature or structurally interesting [ 2 ]. Detailed role definitions and pooled-eligibility rules are provided in Supplementary Methods S1 and Supplementary Table S1 . 2.6. Development, holdout, and non-pooled partitions Development-set material was separated from holdout-eligible material at the level of slice role and row status. Development slices could be used to define comparator logic, test whether the framework behaved sensibly on a known decision problem, and identify where row registration remained incomplete, but development use did not imply pooled eligibility. Holdout-eligible material required a cleaner evidential state, including defined row identities, bounded site claims, and comparator logic strong enough to support later cross-slice interpretation. Slices retained as calibration, controlled buildout, or falsification material were therefore kept outside pooled confusion metrics unless they were later promoted under an explicitly justified evidential rule set. This separation was adopted because cryptic-site benchmarking is especially vulnerable to apparent performance gains that arise from mixing immature and mature cases rather than from genuine discrimination [ 2 ]. 2.7. Frozen baseline definition A static comparator layer was defined once and then frozen for each slice. The purpose of the baseline was not to solve the full cryptic-pocket problem, but to provide a stable reference against which framework decisions could later be compared on the same evaluable rows. For candidate systems requiring public structural support, baseline selection relied on deposited structures and primary source interpretation rather than retrospective convenience [ 4 , 5 ]. In slice-expansion work, this meant that ligand-free or reference structures had to be chosen early and then held fixed during later curation, even when more complex bound-state or complex-context structures were also available. This rule was particularly important for kinase and allosteric systems, where multiple structural states can be found in public databases and where post hoc baseline substitution could otherwise shift the benchmark task itself [ 4 – 7 ]. Once locked, the baseline state for a slice was not allowed to drift during downstream interpretation. 2.8. Frozen decision rules The framework was applied under fixed decision logic. The locked rule set included the chosen observables, threshold logic, replicate handling, missing-data rules, label mapping, and a no-retuning policy. These rules were intended to prevent rescue by reinterpretation after difficult slices or ambiguous rows were encountered. Under this design, weak or unresolved cases were not repaired by narrative argument alone. Instead, they were retained as provisional, calibration-only, restricted-comparator, or non-pooled material until their evidential state changed. The same principle was applied across slice types, including same-series allosteric comparators, mutant cavity anchors, restricted positive-only calibration slices, and falsification cases built around competing-site or occupancy-conflict logic [ 8 – 16 ]. In the present manuscript, the frozen-rule framework functions both as a method and as a boundary condition on what is allowed to count as benchmark evidence. 2.9. Benchmark data model and core tracking tables Benchmark assembly was recorded in a fixed set of core tracking tables. These included candidate_systems.tsv for candidate-slice identity and high-level status, holdout_panel_master.tsv for row-level benchmark registration, baseline_results.tsv for frozen static-comparator outputs, gen2_v1_results.tsv for row-level framework outputs, confusion_metrics_summary.tsv for pooled performance summaries restricted to eligible material, and failure_mode_review.tsv for curated falsification or wrong-hypothesis cases. This table architecture was used to ensure that slice identity, row identity, benchmark role, evaluability, and pooled-metric eligibility were recorded explicitly rather than inferred later from narrative notes. In practical terms, a row was not considered part of benchmark evidence until it existed in the relevant tracking layer with a defined role and an auditable status. 2.10. Slice-level curation and row registration Each slice underwent row-by-row curation before any benchmark role was treated as final. For every candidate system, the curation task was to determine which ligands belonged in the slice, which did not, which rows were evaluable, which remained provisional or unevaluated, and which had to stay outside pooled metrics. This process was deliberately stricter than simple literature collection. Ligands were not admitted merely because they were associated with a target in prior publications. Instead, each row had to map to a specific ligand-target-site hypothesis with a bounded evidence class and a defined benchmark role. Where panel maturity was incomplete, the correct action was not to inflate the slice, but to retain unresolved rows as staged, restricted, calibration-only, or excluded material until the evidential state improved. A full slice-level admissibility matrix for the current panel is provided in Supplementary Table S2 . 2.11. Slice-specific comparator definition: PTP1B For the PTP1B allosteric slice, comparator selection followed explicit inclusion rules. A negative comparator had to be directly tested against PTP1B, explicitly reported as inactive or no-effect, and traceable to a primary source rather than to secondary annotation alone. Under these rules, compound 4 from the Wiesmann benzofuran series was selected as the strongest same-series explicit negative because it was reported as biochemically weak or inactive, with IC50 greater than 500 µM, and as cell-inactive at 500 µM in an insulin-receptor phosphorylation assay [ 8 ]. This made it more suitable for a same-series triage benchmark than secondary negatives from different allosteric mechanism families. MSI-1459 was recorded as a credible secondary inactive analog control, but it was not treated as a same-pocket equivalent because its mechanism family and site logic were less directly aligned with the Wiesmann benzofuran benchmark context [ 9 ]. Accordingly, the PTP1B slice was defined as a minimal same-series allosteric benchmark panel composed of a strong anchor positive, a weaker same-series positive comparator, and an explicit same-series negative control [ 8 , 9 ]. Expanded evidence-tier detail for the PTP1B slice is provided in Supplementary Table S3 . 2.12. Slice-state assignment rules Slice-state assignment followed the benchmark ontology defined above and was based on the relationship between evidential strength, comparator structure, and intended benchmark use. A slice could therefore be retained as a minimal benchmark slice, a limited calibration slice, a limited comparator slice, controlled mature-slice buildout, or falsification material without forcing all systems into pooled evaluation. This rule was particularly important for asymmetric panels in which some targets offered explicit same-series negatives, some offered only bounded positive calibration value, and some were informative mainly as failure-mode tests. Under this design, differences in slice maturity were handled by role restriction rather than by informal narrative adjustment. The same logic also allowed mature expansion candidates, including HIF-2α PAS-B antagonists and MEK1 type III versus ATP-site classification systems, to remain outside the active pooled panel until their slice definitions were sufficiently locked [ 13 – 16 ]. Slice-level admissibility and restricted-use structure for the current panel are summarized in Supplementary Table S2 , with candidate expansion logic listed in Supplementary Table S7 . 2.13. Conditions for pooled scoring and failure-mode handling Pooled scoring was permitted only for rows and slices that satisfied the locked evaluability criteria. A row could contribute to pooled performance summaries only if ligand identity, target context, proposed site, evidence class, and benchmark role were all sufficiently resolved to support cross-slice comparison. Rows with unresolved identity, ambiguous site assignment, incomplete comparator logic, or explicitly restricted slice roles were retained in the benchmark infrastructure but excluded from pooled scoring. This rule was adopted to prevent inflation of benchmark breadth through administrative inclusion of material that had not yet matured into a comparable benchmark unit [ 2 ]. Failure-mode and falsification cases were handled separately from pooled benchmark scoring. Such slices were retained because they tested whether the framework could reject mechanistically contradictory or weakly supported hypotheses under fixed criteria, but they were not counted automatically as standard pooled negative rows. This distinction allowed falsification material to strengthen framework interpretation without allowing it to distort pooled performance summaries or to masquerade as ordinary holdout evidence. 2.14. Statistical and reporting policy This manuscript was written under a primary-endpoint discipline. Each major benchmark claim was tied to one declared decision readout, and supporting metrics were treated as secondary rather than as interchangeable rescue criteria. For benchmark-construction claims, the primary outputs were slice status, row status, benchmark role, and pooled-metric eligibility, rather than descriptive counts of files, structures, or literature mentions. Negative, failed, and indeterminate outcomes were treated as valid outputs of the framework and were reported directly when they altered slice role or blocked pooled interpretation. This reporting policy was chosen to keep the manuscript aligned with the actual evidence ceiling of the current program. Figures and tables were selected on the same principle. Main-text items were intended to carry decisions, not merely implementation detail. Accordingly, the central tables in this study were designed to show slice role, row registration, pooled eligibility, and failure-mode handling, while larger archival or diagnostic material was reserved for supplementary presentation. This policy was intended to keep the manuscript interpretable under reviewer scrutiny and to ensure that the strongest claims remained matched to the strongest evidence. 3. Results 3.1. The current Gen2 panel is governed but not uniformly evaluable The current Gen2 benchmark panel was organized through explicit slice-role assignment rather than through a single pooled benchmark label (Table 1). At the preserved 2026-04-04 checkpoint, no slice met the conditions for active pooled evaluation. PTP1B remained a provenance-aligned but identity-restrained non-row-ready slice, KRAS G12D remained a strengthened strict-live-artifact non-row-ready slice, TP53 Y220C was restricted to calibration-only use, CK2 was retained only as falsification material, and HIF-2α remained policy-closed and outside holdout and pooled metrics. Additional mature-slice candidates, including MEK1 and AKT1 reserve, were not promoted into the active panel. Taken together, these assignments define a governed benchmark panel with explicit slice roles, but without a reviewer-defensible basis for pooled scoring. Table 1. Current Gen2 benchmark panel state and slice-role assignments at the preserved 2026-04-04 checkpoint Table 1. Current Gen2 benchmark panel state and slice-role assignments at the preserved 2026-04-04 checkpoint Slice Biological/task class Current benchmark role Strongest anchor or defining evidence Comparator or falsification status Pooled-metric eligible Current verdict PTP1B Allosteric phosphatase benchmark candidate Provenance-aligned, identity-restrained non-row-ready slice PTP1B_SRC_001 remains the only strong fresh primary anchor source No ligand row-ready; no applied row mapping No Re-parked; not currently evaluable TP53 Y220C Mutant cavity calibration case Calibration-only, explicitly non-evaluable slice Prior mutant-cavity anchor logic retained only for calibration use Not admitted as pooled benchmark material No Restricted to calibration use KRAS G12D Mutant cryptic-pocket anchor case Strengthened strict-live-artifact non-row-ready slice Locked canonical apo/holo pair exists Canonical pair present but row-readiness absent No Re-parked; not currently evaluable CK2 Wrong-hypothesis / competing-site autopsy Excluded falsification material ATP-site occupancy-conflict rejection logic Used only as failure-mode material No Excluded from pooled benchmarking HIF-2α Controlled mature-slice buildout candidate Policy-closed, outside holdout and pooled metrics Mature-slice rationale exists but slice remains policy-closed Not active in current pooled program No Policy-closed MEK1 Candidate mature expansion slice Candidate only, not promoted in current checkpoint External mature-slice candidate logic No live promoted slice No Not part of the active panel AKT1 reserve Reserve expansion candidate Candidate only, not promoted in current checkpoint Reserve mature-slice logic only No live promoted slice No Not part of the active panel Abbreviation: PTP1B_SRC_001, the only strong fresh primary anchor source currently retained for the PTP1B slice. 3.2. PTP1B is the only current minimal benchmark-like slice PTP1B was the only slice in the present program that contained the minimum evidential structure needed for a benchmark-style triage question (Table 2). Within the curated same-series panel, FRJ served as the anchor positive, BB3 as the weaker comparator, and compound 4 as the explicit tested negative. This gave the slice a bounded go / weaker-go / no-go structure rather than a positive-only literature set. The resulting question was therefore not whether PTP1B could support broad allosteric benchmarking, but whether Gen2 could distinguish a clearly supported same-series allosteric ligand from a weaker comparator and from a same-series compound that should be rejected. Under that definition, PTP1B was the only current slice that most closely approximated a minimal benchmark-like unit, although it remained bounded in scope and should not be interpreted as a broad multi-scaffold benchmark. Expanded PTP1B evidence-tier detail is provided in Supplementary Table S3 . Table 2. PTP1B minimal same-series benchmark structure and bounded interpretation Ligand Role in slice Evidence status Benchmark function Key limitation FRJ Anchor positive Strong positive anchor Represents the clearly supported allosteric binding hypothesis Same-series context only BB3 Weaker positive comparator Positive with weaker support than FRJ Tests whether Gen2 preserves graded discrimination within the same series Not intended as a co-equal anchor Compound 4 Explicit same-series negative Explicitly tested weak or inactive comparator Tests rejection of a same-series compound that should not be advanced No direct structural confirmation of site binding Bounded interpretation: PTP1B is the only current slice that behaves like a real benchmark unit rather than restricted support material. 3.3. TP53 Y220C is admissible only as a limited positive-only calibration slice TP53 Y220C was not promoted to pooled benchmark status because the current evidence supported only a restricted positive-only calibration role (Table 3). Within the curated slice, PhiKan083 served as the anchor positive, PK7088 was retained only as a weaker comparator under explicit caution, and MB710 was excluded. No explicit negative comparator was available. The resulting slice therefore supported a bounded calibration question rather than a full benchmark discrimination task. Under this definition, TP53 Y220C was admissible only as restricted calibration material and not as pooled-evaluable benchmark evidence. Expanded TP53 Y220C evidence-tier detail is provided in Supplementary Table S4 and illustrated schematically in Supplementary Figure S1 . Table 3. TP53 Y220C limited positive-only calibration structure and bounded interpretation Table 3. TP53 Y220C slice-role restriction and bounded benchmark use Anchor positive Weaker comparator Excluded candidate Explicit negative Allowed role Pooled-metric eligible Verdict PhiKan083 PK7088, with caution MB710 None Limited positive-only calibration slice No Admissible only as restricted calibration material 3.4. KRAS G12D remains a restricted anchor-plus-weaker-comparator slice KRAS G12D was not promoted to pooled benchmark status because the current evidence supported only a restricted anchor-plus-weaker-comparator role (Table 4). Within the curated slice, MRTX1133 served as the strong anchor, TH-Z835 was retained only as a weaker comparator under explicit caution, and no third ligand was forced into the slice in the absence of adequate support. No explicit negative comparator was available. The resulting slice therefore supported a bounded anchor-versus-weaker-comparator question rather than a benchmark-complete discrimination task. Under this definition, KRAS G12D was admissible only as restricted non-pooled comparator material and not as pooled-evaluable benchmark evidence. Expanded KRAS G12D evidence-tier detail is provided in Supplementary Table S5 and illustrated schematically in Supplementary Figure S2 . Table 4. KRAS G12D restricted anchor-plus-weaker-comparator structure and bounded interpretation Table 4. KRAS G12D restricted slice structure and bounded interpretation Slice Anchor positive Weaker comparator Third ligand / explicit negative Allowed role Bounded verdict KRAS G12D MRTX1133 TH-Z835, with caution No third ligand included; no explicit negative Restricted anchor-plus-weaker-comparator slice; not pooled-metric eligible Admissible only as restricted comparator material 3.5. CK2 is informative only as a falsification slice CK2 was retained not as pooled benchmark evidence, but as a falsification slice that tested rejection behavior under mechanistic contradiction (Table 5). In this case, the value of the slice did not lie in providing another positive or comparator panel. Instead, its value lay in showing that a plausible non-canonical binding hypothesis could still fail under fixed criteria when competing-site or occupancy-conflict logic was taken seriously. CK2 therefore contributed failure-mode coverage rather than benchmark breadth. Under this interpretation, the slice is scientifically useful because it shows that the framework does not only accept plausible positives; it also formalizes when a structurally attractive story should be rejected. Expanded falsification logic for the CK2 slice is provided in Supplementary Table S6 . Table 5. CK2 falsification structure and bounded interpretation Table 5. CK2 falsification-slice structure and bounded interpretation Slice Apparent hypothesis Mechanistic conflict Allowed role Pooled-metric eligible Verdict CK2 Plausible non-canonical or allosteric-like binding story ATP-site occupancy-conflict / competing-site contradiction Falsification slice only No Informative only as rejection and failure-mode material 3.6. The current panel blocks pooled benchmark claims by design Under the current benchmark rules, the panel does not justify pooled scoring (Figure 1). This is not a missing downstream analysis step, but the direct consequence of slice-level role assignment, row-level evaluability restriction, and exclusion of non-comparable material. PTP1B remained non-row-ready because ligand identity and applied row mapping were still constrained at the checkpoint level, KRAS G12D remained non-row-ready under strengthened strict-live-artifact requirements, TP53 Y220C was restricted to calibration-only use, CK2 was retained only as falsification material, and HIF-2α remained policy-closed and outside holdout and pooled metrics. Taken together, these decisions show that the current Gen2 panel supports governed benchmark construction, but not a reviewer-defensible pooled benchmark claim. Figure 1. Program-level pooled-scoring gate for the current Gen2 panel. The current Gen2 panel contains multiple slice roles, but none meets the conditions required for pooled benchmark use at the preserved 2026-04-04 checkpoint. PTP1B and KRAS G12D remain non-row-ready, TP53 Y220C is calibration-only, CK2 is falsification material, and HIF-2α is policy-closed. The resulting program state therefore supports governed benchmark construction but blocks pooled scoring by design. 3.7. Cross-case synthesis Across the current panel, Gen2 does not behave as a simple accept-or-reject workflow. Instead, it separates candidate systems into bounded evidential roles according to what each slice can legitimately support (Table 6). In the present build, PTP1B provides the closest approach to a benchmark-like unit because it contains a same-series positive anchor, a weaker comparator, and an explicit negative, even though pooled promotion remains constrained at the preserved checkpoint. TP53 Y220C is admissible only as a limited positive-only calibration slice, KRAS G12D only as a restricted anchor-plus-weaker-comparator slice, and CK2 only as a falsification slice that tests rejection behavior under mechanistic contradiction. HIF-2α remains policy-closed, while MEK1 and AKT1 reserve remain expansion candidates rather than active benchmark slices. The cross-case result is therefore not pooled benchmark performance, but a framework-level distinction between benchmark-like, calibration, restricted-comparator, falsification, and non-promoted material. That distinction is the main current benchmark result of Gen2. An expanded admissibility summary for the full panel is provided in Supplementary Table S2 . Table 6. Cross-case synthesis of current Gen2 slice roles and admissible uses Table 6. Cross-case synthesis of current Gen2 slice roles and admissible uses Slice Role Anchor or defining evidence Comparator structure Negative or falsification logic Pooled eligible Cross-case verdict PTP1B Closest current benchmark-like slice FRJ anchor positive; same-series allosteric support BB3 as weaker same-series comparator Compound 4 as explicit same-series negative No Closest current approximation to a minimal benchmark unit, but not pooled-promoted at the preserved checkpoint TP53 Y220C Limited calibration slice PhiKan083 as anchor positive PK7088 as weaker comparator with caution No explicit negative; MB710 excluded No Admissible only as positive-only calibration material KRAS G12D Restricted comparator slice MRTX1133 as strong anchor TH-Z835 as weaker comparator with caution No explicit negative; no third ligand included No Admissible only as restricted anchor-plus-weaker-comparator material CK2 Falsification slice ATP-site occupancy-conflict rejection logic Not used as a graded positive-comparator slice Mechanistic contradiction / competing-site conflict No Informative only as a rejection and failure-mode case HIF-2α Policy-closed mature-slice buildout Mature cavity rationale with locked anchor logic Comparator logic exists but not active in pooled panel Excluded by policy closure No Outside holdout and pooled metrics at the current checkpoint MEK1 Candidate mature expansion slice Clean type III allosteric versus ATP-site benchmark potential Comparator logic externally defined Not promoted into active panel No Expansion candidate only AKT1 reserve Reserve expansion candidate Reserve allosteric benchmark potential Baseline and comparator logic not yet promoted Not promoted into active panel No Reserve candidate only Cross-case verdicts are role-based interpretations at the preserved 2026-04-04 checkpoint and do not imply pooled benchmark eligibility. 4. Discussion 4.1. What Gen2 demonstrates at the current checkpoint To our knowledge, Gen2 is an early benchmark-governance framework for cryptic-pocket triage that formalizes slice-level admissibility, restricted-use roles, and pooled-scoring exclusion rules rather than treating all curated cases as uniformly benchmarkable. At the current checkpoint, Gen2 does not demonstrate pooled benchmark performance. What it demonstrates instead is something narrower but methodologically important: benchmark construction in cryptic-pocket discovery can be governed by explicit evidential rules rather than by retrospective inclusion of whatever cases appear most attractive. Under this framework, slices are not treated as automatically equivalent simply because they involve related targets or plausible ligand-pocket stories. They are separated according to what the available evidence can actually support. In the present panel, that meant retaining one slice as the closest current approximation to a minimal benchmark-like unit, restricting two others to bounded non-pooled roles, retaining one as falsification material, and keeping additional candidates outside the active panel. The main result is therefore not a benchmark score, but a disciplined demonstration that Gen2 can prevent weakly comparable, incompletely defined, or mechanistically contradictory material from being counted as benchmark evidence before reviewer-defensible conditions are met. 4.2. Why blocked pooled scoring is a strength, not a failure A weak benchmark can be more damaging than no pooled benchmark at all. In cryptic-pocket discovery, the temptation to combine heterogeneous cases into a single apparent performance summary is strong, particularly when some systems carry plausible ligand-pocket stories, some have partial structural support, and some have only restricted comparator logic. The present study took the opposite approach. Instead of forcing every curated slice into pooled evaluation, Gen2 treated non-comparability itself as an actionable result. Under that rule, blocked pooled scoring is not evidence of framework failure. It is evidence that the admission criteria are functioning as intended. This distinction matters because a pooled benchmark number is only meaningful if the underlying rows are genuinely comparable. If non-row-ready slices, calibration-only material, policy-closed systems, and falsification cases are allowed to enter the same pooled analysis, the resulting score may look quantitative while resting on a mixed evidential base. Such a result would be more vulnerable to reviewer criticism than a deliberately blocked benchmark, because it would confuse administrative inclusion with evaluable evidence. By keeping those categories separate, the current framework preserves a clear boundary between what can support a benchmark claim and what can only support restricted interpretation. The current checkpoint therefore illustrates a conservative but scientifically preferable outcome. The present checkpoint should be interpreted as insufficient available evidence for pooled benchmarking, rather than as failure of the Gen2 framework. Gen2 does not reward apparent panel breadth at the expense of evidential discipline. Instead, it treats refusal to over-score the panel as part of benchmark quality control. In practical terms, this means that the present manuscript contributes a benchmark framework that is willing to stop before pooled evaluation when the slice set does not justify it. That restraint is methodologically stronger than reporting an attractive benchmark number built from cases that were not ready to be compared on equal terms. 4.3. What the current panel still lacks and how it can be strengthened The present panel is structured enough to support a benchmark-construction paper, but it is not yet broad enough to support a mature pooled benchmark claim. The main limitation is not simply the number of targets. It is the incomplete coverage of benchmark functions. At the current checkpoint, the panel contains one slice that most closely approaches a benchmark-like triage unit, two slices that are restricted to bounded non-pooled roles, one falsification slice, and no mature pooled-evaluable set. This means that the program still lacks a sufficiently diverse collection of row-ready slices with clean comparator structure, explicit admissibility logic, and comparable evidential maturity. The most useful next strengthening step is therefore not to add more targets indiscriminately, but to add slices that fill missing benchmark functions. In this context, mature buildout candidates such as HIF-2α and MEK1 are valuable because they address gaps that are not covered by the current panel. HIF-2α offers a non-kinase internal-cavity case with stronger mature-slice potential, while MEK1 offers a cleaner type III versus ATP-site discrimination problem. By contrast, reserve candidates should remain outside the active panel until their baseline state, comparator logic, and slice definitions are fixed to the same standard. Under this logic, benchmark expansion should be driven by missing evidential function rather than by target novelty alone. Candidate mature-slice expansion logic is summarized in Supplementary Table S7 . A second strengthening route is internal rather than expansive. The current panel would also become materially stronger if restricted slices were upgraded to row-ready status under the same fixed rules already used here. In practice, that means resolving the remaining constraints that block pooled use, rather than relaxing the admission threshold. This is especially important because the current paper is stronger when it shows that the same evidential standards apply both to slice admission and to later attempts at panel expansion. The benchmark therefore does not need arbitrary growth. It needs either additional mature slices that fill missing benchmark roles, or stricter maturation of the slices already present. 4.4. Limitations Several limitations define the evidential ceiling of the present study. First, the current panel does not support pooled benchmark claims. That restriction is intentional, but it also means that the manuscript cannot yet provide a cross-slice performance summary, comparative accuracy statistic, or broad generality claim for Gen2. The strongest contribution at this stage is therefore framework-level governance rather than completed benchmark performance. This should be interpreted as insufficient available evidence for pooled benchmarking, rather than as failure of the framework itself. Second, the panel remains asymmetric in slice maturity and evidential structure. PTP1B comes closest to a benchmark-like unit because it contains a same-series anchor positive, a weaker comparator, and an explicit negative, but even that slice remains constrained at the preserved checkpoint. TP53 Y220C and KRAS G12D were retained only in restricted non-pooled roles, while CK2 was informative only as falsification material. This asymmetry is scientifically manageable, but it limits how far cross-case comparison can be pushed. Third, not all restricted slices fail for the same reason. Some remain limited because of missing explicit negatives, some because of row-readiness constraints, some because of unresolved mapping or comparator structure, and some because their value lies mainly in rejection behavior rather than benchmark inclusion. As a result, the current study should not be read as showing that all non-pooled slices are equally immature or equally informative. Their restricted status reflects different evidential bottlenecks. Fourth, the benchmark remains dependent on literature-supported and structure-supported slice definitions rather than on a fully matured holdout panel. This is appropriate for a benchmark-construction paper, but it means that some conclusions are still conditional on future slice maturation under the same fixed rules. In particular, the present study does not yet show how Gen2 behaves across a larger set of row-ready slices spanning multiple target classes under a common pooled scoring framework. Finally, the current manuscript does not integrate blinded orthogonal reanalysis or experimental assay follow-up into the benchmark core. That boundary was deliberate, because mixing benchmark construction with validation and assay layers at this stage would have weakened interpretability. However, it also means that the present paper addresses benchmark admissibility rather than full end-to-end validation. Future work will therefore need to test whether the same governance logic remains robust when additional mature slices are promoted and when benchmark-ready slices are challenged under external validation conditions. 4.5. Conclusion This study presents Gen2 not as a completed pooled benchmark, but as a governed benchmark-construction framework for binding hypothesis triage in cryptic-pocket discovery. The main result is that benchmark assembly itself can be treated as a scientific decision layer, with explicit rules for slice admission, role assignment, row registration, pooled-metric eligibility, and falsification handling. Under those rules, the current panel did not justify pooled scoring. Rather than weakening the framework, that outcome showed that Gen2 can prevent weakly comparable, incompletely defined, or mechanistically contradictory cases from being counted as benchmark evidence before reviewer-defensible conditions are met. At the present checkpoint, Gen2 supports one slice that most closely approaches a minimal benchmark-like unit, two slices retained only in restricted non-pooled roles, one falsification slice, and additional mature-slice candidates that remain outside the active panel. Taken together, these assignments define the current contribution of the framework: not a benchmark number, but a reproducible evidential architecture for deciding what may and may not count in a cryptic-pocket benchmark. That is the central conclusion of the present work. The next stage is clear. Gen2 will become a stronger benchmark only if additional slices are matured under the same fixed rules, or if currently restricted slices become row-ready without relaxing admission standards. Until then, the appropriate claim is bounded but meaningful: Gen2 provides a reviewer-defensible framework for benchmark construction in cryptic-pocket discovery, and its present value lies in governed admissibility rather than in pooled performance. Declarations 5. Acknowledgements Not applicable. 6. Funding This research received no external funding. 7. Author Contributions Hoosdally Shakeel conceived the study, designed the benchmark framework, curated the benchmark slices, performed the analysis, interpreted the results, prepared the figures and tables, and wrote the manuscript. 8. Conflicts of Interest The author declares a patent-related competing interest. Intellectual property related to aspects of the framework described in this manuscript has been filed through the USPTO. 9. Data Availability The data supporting the findings of this study are available from the author on reasonable request. These materials include the benchmark tables, slice-role assignments, row-registration outputs, figure source files, and supporting benchmark-construction records needed to interpret the results reported in this manuscript. Because aspects of the framework are the subject of patent-related protection, full release of all underlying workflow materials, scripts, and intermediate development records is not included in the present manuscript. Any shared materials will therefore be limited to those necessary to support transparency, interpretation, and reproducibility of the reported benchmark state without disclosing protected implementation details. References Vajda S, Beglov D, Wakefield AE, Egbert M, Whitty A (2018) Cryptic binding sites on proteins: definition, detection, and druggability. Curr Opin Chem Biol 44:1–8. 10.1016/j.cbpa.2018.05.003 Škrhák V, Novotný M, Feidakis CP, Krivák R, Hoksza D (2025) CryptoBench: cryptic protein-ligand binding sites dataset and benchmark. Bioinformatics 41(1):btae745. 10.1093/bioinformatics/btae745 Shakeel H (2026) Gen2: A Biophysical Triage Framework for Binding Hypotheses in Cryptic Pocket Discovery. 10.26434/chemrxiv.15001255/v1 . ChemRxiv Burley SK, Bhikadiya C, Bi C, Bittrich S, Chen L, Crichlow GV et al (2023) RCSB Protein Data Bank (RCSB.org): delivery of experimentally determined PDB structures alongside one million computed structure models of proteins from artificial intelligence/machine learning. Nucleic Acids Res 51(D1):D488–D508. 10.1093/nar/gkac1077 Sayers EW, Beck J, Bolton EE, Bourexis D, Brister JR, Canese K et al (2025) Database resources of the National Center for Biotechnology Information. Nucleic Acids Res 53(D1):D20–D32. 10.1093/nar/gkae988 Kooistra AJ, Kanev GK, van Linden OPGJ, Leurs R, de Esch IJP, de Graaf C (2016) KLIFS: a structural kinase-ligand interaction database. Nucleic Acids Res 44(D1):D365–D371. 10.1093/nar/gkv1082 He J, Liu X, Zhu C, Zha J, Ou W, Zhang G et al (2024) ASD2023: towards the integrating landscapes of allosteric knowledgebase. Nucleic Acids Res 52(D1):D376–D383. 10.1093/nar/gkad915 Wiesmann C, Barr KJ, Kung J, Zhu J, Erlanson DA, Shen W et al (2004) Allosteric inhibition of protein tyrosine phosphatase 1B. Nat Struct Mol Biol 11(8):730–737. 10.1038/nsmb803 Krishnan N, Koveal D, Miller DH, Xue B, Akshinthala SD, Kragelj J et al (2014) Targeting the disordered C terminus of PTP1B with an allosteric inhibitor. Nat Chem Biol 10(7):558–566. 10.1038/nchembio.1528 Boeckler FM, Joerger AC, Jaggi G, Rutherford TJ, Veprintsev DB, Fersht AR (2008) Targeted rescue of a destabilized mutant of p53 by an in silico screened drug. Proc Natl Acad Sci U S A 105(30):10360–10365. 10.1073/pnas.0805326105 Wang X, Allen S, Blake JF, Bowcut V, Briere DM, Calinisan A et al (2022) Identification of MRTX1133, a noncovalent, potent, and selective KRASG12D inhibitor. J Med Chem 65(4):3123–3133. 10.1021/acs.jmedchem.1c01688 Mao Z, Xiao H, Shen P, Yang Y, Guan Y, Wang Y et al (2022) KRAS(G12D) can be targeted by potent inhibitors via formation of a salt bridge. Cell Discov 8:5. 10.1038/s41421-021-00368-w Wallace EM, Rizzi JP, Han G, Wehn PM, Cao Z, Du X et al (2016) A small-molecule antagonist of HIF2α is efficacious in preclinical models of renal cell carcinoma. Cancer Res 76(18):5491–5500. 10.1158/0008-5472.CAN-16-0473 Ren X, Wang J, Cui S, Gao B, Li M, Sun Y et al (2022) Structural basis for the allosteric inhibition of hypoxia-inducible factor HIF-2 by belzutifan. Mol Pharmacol 102(6):366–376. 10.1124/molpharm.122.000525 Hatzivassiliou G, Haling JR, Chen H, Song K, Price S, Heald R et al (2013) Mechanism of MEK inhibition determines efficacy in mutant KRAS- versus BRAF-driven cancers. Nature 501(7466):232–236. 10.1038/nature12441 Takano K, Munehira Y, Itou J, Watanabe N, Horiguchi M, Takikawa S et al (2023) Discovery of a novel ATP-competitive MEK inhibitor DS03090629 that overcomes resistance conferred by BRAF overexpression in BRAF-mutated melanoma. Mol Cancer Ther 22(3):317–329. 10.1158/1535-7163.MCT-22-0306 Additional Declarations The authors declare no competing interests. Supplementary Files Gen2SupplementaryInformation.docx Supplementary Information Supplementary Information is available for this article and includes Supplementary Methods S1, Supplementary Tables S1-S7, and Supplementary Figures S1-S2. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9336311","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Method Article","associatedPublications":[],"authors":[{"id":618397006,"identity":"a895dc6e-8f27-4389-af97-2579766f0b61","order_by":0,"name":"Mr shakeel Hoosdally","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABD0lEQVRIiWNgGAWjYDACCQZmIGmRAGEXWDDwg1gJBQS1SEC1GEgwSDaAtBiQosXgAIiJR4v87ObHBh8qJPL42dsf3vhhIJG4+fzqxA8PDBjk+cUOYNVicOeYceKMMxLFkj1njC17gFq23Xi7WQLoMMOZsxOwa5FIMD7M2yaRuOFGDpsED1jL2Q0gLQkGt7FrkZ+R/vnw338gLenPJP+AHDbj7OYf+LQw3MgxTmZsAGlJMJMG2bKBv3cbXlsMbuQUG/Yck0icCfSLtYyBhPGMG7zbLBIMJHD6BeiwzRI/amwS+4EhdvNNhY1sf//ZzTd/VNjI80vjcBg6cGyAxRHRwJ6B/wDxqkfBKBgFo2BEAADwz2LU2h+jYAAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0009-0008-4144-718X","institution":"ministry of education and higher scientific research","correspondingAuthor":true,"prefix":"Mr","firstName":"shakeel","middleName":"","lastName":"Hoosdally","suffix":""}],"badges":[],"createdAt":"2026-04-06 17:43:17","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true},"doi":"10.21203/rs.3.rs-9336311/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9336311/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106724967,"identity":"2e584aa8-dc1b-445d-ae99-995ea0cffcdd","added_by":"auto","created_at":"2026-04-12 18:30:48","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":375256,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eProgram-level pooled-scoring gate for the current Gen2 panel.\u003c/strong\u003e The current Gen2 panel contains multiple slice roles, but none meets the conditions required for pooled benchmark use at the preserved 2026-04-04 checkpoint. PTP1B and KRAS G12D remain non-row-ready, TP53 Y220C is calibration-only, CK2 is falsification material, and HIF-2α is policy-closed. The resulting program state therefore supports governed benchmark construction but blocks pooled scoring by design.\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-9336311/v1/179018cfd35d1bbe2ccd437f.jpg"},{"id":106994023,"identity":"ef2d1481-16d7-46d5-bcd5-60c39da6692f","added_by":"auto","created_at":"2026-04-15 15:02:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2211451,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9336311/v1/7efeee56-f319-4054-8bf9-e9c94ef662ec.pdf"},{"id":106540939,"identity":"e2b60919-393e-4815-a3fb-cc138d0bb488","added_by":"auto","created_at":"2026-04-09 16:03:10","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":530157,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary Information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSupplementary Information is available for this article and includes Supplementary Methods S1, Supplementary Tables S1-S7, and Supplementary Figures S1-S2.\u003c/p\u003e","description":"","filename":"Gen2SupplementaryInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-9336311/v1/e21378abe86a1bef1181195b.docx"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eGen2: Building a Reviewer-Defensible Benchmark for Binding Hypothesis Triage in Cryptic Pocket Discovery\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"Simple Summary","content":"\u003cp\u003eBenchmark papers in cryptic-pocket discovery can become misleading when weak, ambiguous, or incomplete cases are forced into pooled scoring. In this study, Gen2 is presented as a benchmark framework that first decides which cases are actually fit for evaluation. Applied to the current panel, this process showed that several candidate slices should remain parked, calibration-only, or excluded rather than being scored prematurely. The main result is therefore not benchmark performance, but a controlled framework that prevents invalid pooled claims and defines what must be resolved before a full benchmark can be done.\u003c/p\u003e"},{"header":"1. Introduction","content":"\u003cp\u003eCryptic-pocket discovery remains methodologically difficult because the central object of interest is often not fully visible in the ligand-free structure. In many systems, a site that later accommodates a ligand is absent, occluded, or too weakly formed to be recognized confidently from a single apo conformation alone [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. This has made cryptic pockets attractive in drug discovery, particularly for targets that appear poorly tractable in their ground-state structures, but it has also created a recurring evidential problem. Methods based on molecular dynamics, mixed-solvent sampling, or other ensemble-based strategies can reveal conformations that static analysis misses, yet they can also generate many transient cavities that are not credible small-molecule binding hypotheses [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The practical difficulty is therefore not only detecting pockets, but deciding which proposed ligand-pocket hypotheses are strong enough to count as benchmark evidence.\u003c/p\u003e \u003cp\u003eThat problem becomes sharper when benchmark claims are made. Recent work on cryptic-site datasets has highlighted a basic weakness in the field: performance can look artificially strong when methods are trained or evaluated mainly on ligand-bound states rather than on matched apo-holo settings, even though cryptic-site prediction is supposed to address precisely the gap between those states [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. In other words, a benchmark can fail before any scoring begins if heterogeneous cases are pooled despite ambiguous labels, unstable site assignment, missing row-level outputs, or mismatched evidential standards. For cryptic-pocket and allosteric systems, this is not a minor bookkeeping issue. It directly affects whether pooled scoring reflects a real discrimination task or only a mixture of partially comparable cases.\u003c/p\u003e \u003cp\u003eWe recently introduced \u003cb\u003eGen2\u003c/b\u003e as a biophysical triage framework for evaluating ligand-pocket hypotheses in cryptic-pocket discovery [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. That earlier study focused on case-level adjudication: whether a proposed binding hypothesis remained credible after stricter biophysical scrutiny in individual systems. The present manuscript addresses a different question. Here, the focus is the benchmark-construction layer itself: how candidate slices are screened, how their roles are bounded, when they can be treated as evaluable, and when they must remain calibration-only, parked, policy-closed, or excluded from pooled scoring. Accordingly, this paper does not present a completed pooled benchmark and does not fold validation or assay layers into the benchmark core. Instead, it asks a narrower question: can Gen2 impose enough evidential discipline to distinguish reviewer-defensible benchmark material from non-evaluable material before pooled scoring is attempted? In this framework, benchmark curation, slice locking, table repair, and explicit marking of unevaluated rows are not background administration. They are part of the scientific result.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Benchmark claim and scope\u003c/h2\u003e \u003cp\u003eThis study was designed as a benchmark-construction analysis rather than as a completed pooled benchmark or an experimental validation study. The purpose was to define how candidate slices enter the benchmark program, how their roles are assigned, and under what conditions they can be treated as evaluable material for later cross-slice comparison. Accordingly, the present Methods focus on benchmark curation, slice locking, row-status assignment, and pooled-metric eligibility. Validation layers such as blinded orthogonal reanalysis, biophysical assay follow-up, and assay-integration logic were kept outside the benchmark core by design. This separation was adopted because cryptic-site benchmarking is especially vulnerable to inflated performance when heterogeneous cases are pooled before their evidential status is made comparable [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Prediction unit\u003c/h2\u003e \u003cp\u003eThe basic unit of analysis was defined as one ligand-target-site hypothesis per row. Each row therefore represented a specific claim: that a named ligand, in a defined target system and structural context, supported or failed to support a defined binding-site hypothesis. This row-level design was chosen to prevent case-level narrative drift and to force explicit treatment of ligand identity, target state, proposed site, evidence class, and benchmark role before any result was counted. In practice, benchmark assembly could not proceed by describing a target in general terms alone. Instead, each proposed benchmark element had to be reducible to auditable row entries that could later be classified as evaluable, unevaluated, excluded, calibration-only, comparator-only, or otherwise restricted.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Ground-truth and label rules\u003c/h2\u003e \u003cp\u003eRow labels were assigned under bounded evidential rules. A positive row required direct or closely linked evidence supporting the stated ligand-target-site hypothesis. A negative row required explicit inactivity, no-effect behavior, or equivalent rejection evidence traceable to a primary source. Rows lacking sufficient support were not forced into positive or negative classes. Instead, they were retained as provisional, unresolved, or unevaluated until their evidential status could be clarified. This rule was intended to prevent false confidence from partial literature support, incomplete structural context, or weak comparator logic. Untested ligands, ambiguous site assignments, unresolved identities, and rows lacking a credible comparator were therefore retained inside the benchmark infrastructure only as non-pooled material. They could inform later slice development, but they were not allowed to contribute to pooled benchmark claims [\u003cspan additionalcitationids=\"CR9 CR10 CR11\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4. Candidate-system screening and slice admission criteria\u003c/h2\u003e \u003cp\u003eCandidate systems were screened using a structural-first benchmark logic. A slice was considered suitable only if it presented a clear pocket-level adjudication question, at least one credible anchor positive, at least one fair comparator, adequate public structural support, and a task definition that could be handled under the existing decision framework without rewriting the charter. Systems were not promoted simply because they were biologically important or well represented in the literature. They were promoted only when the available evidence could be translated into a bounded benchmark problem with explicit row-level registration. Public structural and literature resources used during screening included the RCSB Protein Data Bank, PubMed, KLIFS, and ASD2023 [\u003cspan additionalcitationids=\"CR5 CR6\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. This screening step also distinguished between benchmark roles. Some systems were suitable as minimal benchmark slices, some only as limited calibration or comparator material, and some only as falsification or failure-mode cases. Systems that lacked clean comparator logic, contained unresolved structural ambiguity, or could not be mapped to row-level adjudication without interpretive inflation were not admitted as pooled benchmark material.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5. Benchmark panel ontology\u003c/h2\u003e \u003cp\u003eBefore any slice was interpreted, each candidate system was assigned a bounded benchmark role. The allowed roles were: included benchmark slice, limited calibration slice, limited comparator slice, controlled mature-slice buildout, falsification or failure-mode slice, and excluded slice. This ontology was introduced to prevent heterogeneous cases from being pooled under a single benchmark label when their evidential status was not actually comparable. In practical terms, it allowed the benchmark to distinguish between slices that could contribute to later scoring, slices that could only provide calibration or restricted comparator value, and slices that were useful only for documenting failure or rejection behavior. The purpose of this role system was not to make the panel appear larger, but to prevent weakly supported or mismatched cases from being promoted into benchmark evidence simply because they were available in the literature or structurally interesting [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eDetailed role definitions and pooled-eligibility rules are provided in \u003cb\u003eSupplementary Methods S1\u003c/b\u003e and \u003cb\u003eSupplementary Table S1\u003c/b\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6. Development, holdout, and non-pooled partitions\u003c/h2\u003e \u003cp\u003eDevelopment-set material was separated from holdout-eligible material at the level of slice role and row status. Development slices could be used to define comparator logic, test whether the framework behaved sensibly on a known decision problem, and identify where row registration remained incomplete, but development use did not imply pooled eligibility. Holdout-eligible material required a cleaner evidential state, including defined row identities, bounded site claims, and comparator logic strong enough to support later cross-slice interpretation. Slices retained as calibration, controlled buildout, or falsification material were therefore kept outside pooled confusion metrics unless they were later promoted under an explicitly justified evidential rule set. This separation was adopted because cryptic-site benchmarking is especially vulnerable to apparent performance gains that arise from mixing immature and mature cases rather than from genuine discrimination [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.7. Frozen baseline definition\u003c/h2\u003e \u003cp\u003eA static comparator layer was defined once and then frozen for each slice. The purpose of the baseline was not to solve the full cryptic-pocket problem, but to provide a stable reference against which framework decisions could later be compared on the same evaluable rows. For candidate systems requiring public structural support, baseline selection relied on deposited structures and primary source interpretation rather than retrospective convenience [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. In slice-expansion work, this meant that ligand-free or reference structures had to be chosen early and then held fixed during later curation, even when more complex bound-state or complex-context structures were also available. This rule was particularly important for kinase and allosteric systems, where multiple structural states can be found in public databases and where post hoc baseline substitution could otherwise shift the benchmark task itself [\u003cspan additionalcitationids=\"CR5 CR6\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Once locked, the baseline state for a slice was not allowed to drift during downstream interpretation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e2.8. Frozen decision rules\u003c/h2\u003e \u003cp\u003eThe framework was applied under fixed decision logic. The locked rule set included the chosen observables, threshold logic, replicate handling, missing-data rules, label mapping, and a no-retuning policy. These rules were intended to prevent rescue by reinterpretation after difficult slices or ambiguous rows were encountered. Under this design, weak or unresolved cases were not repaired by narrative argument alone. Instead, they were retained as provisional, calibration-only, restricted-comparator, or non-pooled material until their evidential state changed. The same principle was applied across slice types, including same-series allosteric comparators, mutant cavity anchors, restricted positive-only calibration slices, and falsification cases built around competing-site or occupancy-conflict logic [\u003cspan additionalcitationids=\"CR9 CR10 CR11 CR12 CR13 CR14 CR15\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. In the present manuscript, the frozen-rule framework functions both as a method and as a boundary condition on what is allowed to count as benchmark evidence.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e2.9. Benchmark data model and core tracking tables\u003c/h2\u003e \u003cp\u003eBenchmark assembly was recorded in a fixed set of core tracking tables. These included \u003cspan fontcategory=\"NonProportional\" class=\"\" name=\"Emphasis\"\u003ecandidate_systems.tsv\u003c/span\u003e for candidate-slice identity and high-level status, \u003cspan fontcategory=\"NonProportional\" class=\"\" name=\"Emphasis\"\u003eholdout_panel_master.tsv\u003c/span\u003e for row-level benchmark registration, \u003cspan fontcategory=\"NonProportional\" class=\"\" name=\"Emphasis\"\u003ebaseline_results.tsv\u003c/span\u003e for frozen static-comparator outputs, \u003cspan fontcategory=\"NonProportional\" class=\"\" name=\"Emphasis\"\u003egen2_v1_results.tsv\u003c/span\u003e for row-level framework outputs, \u003cspan fontcategory=\"NonProportional\" class=\"\" name=\"Emphasis\"\u003econfusion_metrics_summary.tsv\u003c/span\u003e for pooled performance summaries restricted to eligible material, and \u003cspan fontcategory=\"NonProportional\" class=\"\" name=\"Emphasis\"\u003efailure_mode_review.tsv\u003c/span\u003e for curated falsification or wrong-hypothesis cases. This table architecture was used to ensure that slice identity, row identity, benchmark role, evaluability, and pooled-metric eligibility were recorded explicitly rather than inferred later from narrative notes. In practical terms, a row was not considered part of benchmark evidence until it existed in the relevant tracking layer with a defined role and an auditable status.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e2.10. Slice-level curation and row registration\u003c/h2\u003e \u003cp\u003eEach slice underwent row-by-row curation before any benchmark role was treated as final. For every candidate system, the curation task was to determine which ligands belonged in the slice, which did not, which rows were evaluable, which remained provisional or unevaluated, and which had to stay outside pooled metrics. This process was deliberately stricter than simple literature collection. Ligands were not admitted merely because they were associated with a target in prior publications. Instead, each row had to map to a specific ligand-target-site hypothesis with a bounded evidence class and a defined benchmark role. Where panel maturity was incomplete, the correct action was not to inflate the slice, but to retain unresolved rows as staged, restricted, calibration-only, or excluded material until the evidential state improved.\u003c/p\u003e \u003cp\u003eA full slice-level admissibility matrix for the current panel is provided in \u003cb\u003eSupplementary Table S2\u003c/b\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e2.11. Slice-specific comparator definition: PTP1B\u003c/h2\u003e \u003cp\u003eFor the PTP1B allosteric slice, comparator selection followed explicit inclusion rules. A negative comparator had to be directly tested against PTP1B, explicitly reported as inactive or no-effect, and traceable to a primary source rather than to secondary annotation alone. Under these rules, compound 4 from the Wiesmann benzofuran series was selected as the strongest same-series explicit negative because it was reported as biochemically weak or inactive, with IC50 greater than 500 \u0026micro;M, and as cell-inactive at 500 \u0026micro;M in an insulin-receptor phosphorylation assay [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. This made it more suitable for a same-series triage benchmark than secondary negatives from different allosteric mechanism families. MSI-1459 was recorded as a credible secondary inactive analog control, but it was not treated as a same-pocket equivalent because its mechanism family and site logic were less directly aligned with the Wiesmann benzofuran benchmark context [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Accordingly, the PTP1B slice was defined as a minimal same-series allosteric benchmark panel composed of a strong anchor positive, a weaker same-series positive comparator, and an explicit same-series negative control [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eExpanded evidence-tier detail for the PTP1B slice is provided in \u003cb\u003eSupplementary Table S3\u003c/b\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e2.12. Slice-state assignment rules\u003c/h2\u003e \u003cp\u003eSlice-state assignment followed the benchmark ontology defined above and was based on the relationship between evidential strength, comparator structure, and intended benchmark use. A slice could therefore be retained as a minimal benchmark slice, a limited calibration slice, a limited comparator slice, controlled mature-slice buildout, or falsification material without forcing all systems into pooled evaluation. This rule was particularly important for asymmetric panels in which some targets offered explicit same-series negatives, some offered only bounded positive calibration value, and some were informative mainly as failure-mode tests. Under this design, differences in slice maturity were handled by role restriction rather than by informal narrative adjustment. The same logic also allowed mature expansion candidates, including HIF-2α PAS-B antagonists and MEK1 type III versus ATP-site classification systems, to remain outside the active pooled panel until their slice definitions were sufficiently locked [\u003cspan additionalcitationids=\"CR14 CR15\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSlice-level admissibility and restricted-use structure for the current panel are summarized in \u003cb\u003eSupplementary Table S2\u003c/b\u003e, with candidate expansion logic listed in \u003cb\u003eSupplementary Table S7\u003c/b\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e2.13. Conditions for pooled scoring and failure-mode handling\u003c/h2\u003e \u003cp\u003ePooled scoring was permitted only for rows and slices that satisfied the locked evaluability criteria. A row could contribute to pooled performance summaries only if ligand identity, target context, proposed site, evidence class, and benchmark role were all sufficiently resolved to support cross-slice comparison. Rows with unresolved identity, ambiguous site assignment, incomplete comparator logic, or explicitly restricted slice roles were retained in the benchmark infrastructure but excluded from pooled scoring. This rule was adopted to prevent inflation of benchmark breadth through administrative inclusion of material that had not yet matured into a comparable benchmark unit [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFailure-mode and falsification cases were handled separately from pooled benchmark scoring. Such slices were retained because they tested whether the framework could reject mechanistically contradictory or weakly supported hypotheses under fixed criteria, but they were not counted automatically as standard pooled negative rows. This distinction allowed falsification material to strengthen framework interpretation without allowing it to distort pooled performance summaries or to masquerade as ordinary holdout evidence.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e2.14. Statistical and reporting policy\u003c/h2\u003e \u003cp\u003eThis manuscript was written under a primary-endpoint discipline. Each major benchmark claim was tied to one declared decision readout, and supporting metrics were treated as secondary rather than as interchangeable rescue criteria. For benchmark-construction claims, the primary outputs were slice status, row status, benchmark role, and pooled-metric eligibility, rather than descriptive counts of files, structures, or literature mentions. Negative, failed, and indeterminate outcomes were treated as valid outputs of the framework and were reported directly when they altered slice role or blocked pooled interpretation. This reporting policy was chosen to keep the manuscript aligned with the actual evidence ceiling of the current program.\u003c/p\u003e \u003cp\u003eFigures and tables were selected on the same principle. Main-text items were intended to carry decisions, not merely implementation detail. Accordingly, the central tables in this study were designed to show slice role, row registration, pooled eligibility, and failure-mode handling, while larger archival or diagnostic material was reserved for supplementary presentation. This policy was intended to keep the manuscript interpretable under reviewer scrutiny and to ensure that the strongest claims remained matched to the strongest evidence.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cp\u003e\u003cstrong\u003e3.1. The current Gen2 panel is governed but not uniformly evaluable\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe current Gen2 benchmark panel was organized through explicit slice-role assignment rather than through a single pooled benchmark label (Table 1). At the preserved 2026-04-04 checkpoint, no slice met the conditions for active pooled evaluation. PTP1B remained a provenance-aligned but identity-restrained non-row-ready slice, KRAS G12D remained a strengthened strict-live-artifact non-row-ready slice, TP53 Y220C was restricted to calibration-only use, CK2 was retained only as falsification material, and HIF-2\u0026alpha; remained policy-closed and outside holdout and pooled metrics. Additional mature-slice candidates, including MEK1 and AKT1 reserve, were not promoted into the active panel. Taken together, these assignments define a governed benchmark panel with explicit slice roles, but without a reviewer-defensible basis for pooled scoring.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1. Current Gen2 benchmark panel state and slice-role assignments at the preserved 2026-04-04 checkpoint\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1. Current Gen2 benchmark panel state and slice-role assignments at the preserved 2026-04-04 checkpoint\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSlice\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBiological/task class\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCurrent benchmark role\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eStrongest anchor or defining evidence\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eComparator or falsification status\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePooled-metric eligible\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCurrent verdict\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 82px;\"\u003e\n \u003cp\u003ePTP1B\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eAllosteric phosphatase benchmark candidate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 168px;\"\u003e\n \u003cp\u003eProvenance-aligned, identity-restrained non-row-ready slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 178px;\"\u003e\n \u003cp\u003ePTP1B_SRC_001 remains the only strong fresh primary anchor source\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eNo ligand row-ready; no applied row mapping\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eRe-parked; not currently evaluable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 82px;\"\u003e\n \u003cp\u003eTP53 Y220C\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eMutant cavity calibration case\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 168px;\"\u003e\n \u003cp\u003eCalibration-only, explicitly non-evaluable slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 178px;\"\u003e\n \u003cp\u003ePrior mutant-cavity anchor logic retained only for calibration use\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eNot admitted as pooled benchmark material\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eRestricted to calibration use\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 82px;\"\u003e\n \u003cp\u003eKRAS G12D\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eMutant cryptic-pocket anchor case\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 168px;\"\u003e\n \u003cp\u003eStrengthened strict-live-artifact non-row-ready slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 178px;\"\u003e\n \u003cp\u003eLocked canonical apo/holo pair exists\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eCanonical pair present but row-readiness absent\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eRe-parked; not currently evaluable\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 82px;\"\u003e\n \u003cp\u003eCK2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eWrong-hypothesis / competing-site autopsy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 168px;\"\u003e\n \u003cp\u003eExcluded falsification material\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 178px;\"\u003e\n \u003cp\u003eATP-site occupancy-conflict rejection logic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eUsed only as failure-mode material\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eExcluded from pooled benchmarking\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 82px;\"\u003e\n \u003cp\u003eHIF-2\u0026alpha;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eControlled mature-slice buildout candidate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 168px;\"\u003e\n \u003cp\u003ePolicy-closed, outside holdout and pooled metrics\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 178px;\"\u003e\n \u003cp\u003eMature-slice rationale exists but slice remains policy-closed\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eNot active in current pooled program\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003ePolicy-closed\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 82px;\"\u003e\n \u003cp\u003eMEK1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eCandidate mature expansion slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 168px;\"\u003e\n \u003cp\u003eCandidate only, not promoted in current checkpoint\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 178px;\"\u003e\n \u003cp\u003eExternal mature-slice candidate logic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eNo live promoted slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eNot part of the active panel\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 82px;\"\u003e\n \u003cp\u003eAKT1 reserve\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eReserve expansion candidate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 168px;\"\u003e\n \u003cp\u003eCandidate only, not promoted in current checkpoint\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 178px;\"\u003e\n \u003cp\u003eReserve mature-slice logic only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 163px;\"\u003e\n \u003cp\u003eNo live promoted slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 101px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 139px;\"\u003e\n \u003cp\u003eNot part of the active panel\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cem\u003eAbbreviation: PTP1B_SRC_001, the only strong fresh primary anchor source currently retained for the PTP1B slice.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.2. PTP1B is the only current minimal benchmark-like slice\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePTP1B was the only slice in the present program that contained the minimum evidential structure needed for a benchmark-style triage question (Table 2). Within the curated same-series panel, FRJ served as the anchor positive, BB3 as the weaker comparator, and compound 4 as the explicit tested negative. This gave the slice a bounded go / weaker-go / no-go structure rather than a positive-only literature set. The resulting question was therefore not whether PTP1B could support broad allosteric benchmarking, but whether Gen2 could distinguish a clearly supported same-series allosteric ligand from a weaker comparator and from a same-series compound that should be rejected. Under that definition, PTP1B was the only current slice that most closely approximated a minimal benchmark-like unit, although it remained bounded in scope and should not be interpreted as a broad multi-scaffold benchmark. Expanded PTP1B evidence-tier detail is provided in \u003cstrong\u003eSupplementary Table S3\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2. PTP1B minimal same-series benchmark structure and bounded interpretation\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 86px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLigand\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRole in slice\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 149px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eEvidence status\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBenchmark function\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 149px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eKey limitation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eFRJ\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eAnchor positive\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eStrong positive anchor\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eRepresents the clearly supported allosteric binding hypothesis\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eSame-series context only\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eBB3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eWeaker positive comparator\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003ePositive with weaker support than FRJ\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eTests whether Gen2 preserves graded discrimination within the same series\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eNot intended as a co-equal anchor\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eCompound 4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eExplicit same-series negative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eExplicitly tested weak or inactive comparator\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eTests rejection of a same-series compound that should not be advanced\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eNo direct structural confirmation of site binding\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cem\u003eBounded interpretation: PTP1B is the only current slice that behaves like a real benchmark unit rather than restricted support material.\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.3. TP53 Y220C is admissible only as a limited positive-only calibration slice\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTP53 Y220C was not promoted to pooled benchmark status because the current evidence supported only a restricted positive-only calibration role (Table 3). Within the curated slice, PhiKan083 served as the anchor positive, PK7088 was retained only as a weaker comparator under explicit caution, and MB710 was excluded. No explicit negative comparator was available. The resulting slice therefore supported a bounded calibration question rather than a full benchmark discrimination task. Under this definition, TP53 Y220C was admissible only as restricted calibration material and not as pooled-evaluable benchmark evidence. Expanded TP53 Y220C evidence-tier detail is provided in \u003cstrong\u003eSupplementary Table S4\u003c/strong\u003e and illustrated schematically in \u003cstrong\u003eSupplementary Figure S1\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3. TP53 Y220C limited positive-only calibration structure and bounded interpretation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3. TP53 Y220C slice-role restriction and bounded benchmark use\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 91px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAnchor positive\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 110px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eWeaker comparator\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 91px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eExcluded candidate\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 77px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eExplicit negative\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAllowed role\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePooled-metric eligible\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 154px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eVerdict\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 91px;\"\u003e\n \u003cp\u003ePhiKan083\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 110px;\"\u003e\n \u003cp\u003ePK7088, with caution\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 91px;\"\u003e\n \u003cp\u003eMB710\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 77px;\"\u003e\n \u003cp\u003eNone\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eLimited positive-only calibration slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 96px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 154px;\"\u003e\n \u003cp\u003eAdmissible only as restricted calibration material\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cstrong\u003e3.4. KRAS G12D remains a restricted anchor-plus-weaker-comparator slice\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eKRAS G12D was not promoted to pooled benchmark status because the current evidence supported only a restricted anchor-plus-weaker-comparator role (Table 4). Within the curated slice, MRTX1133 served as the strong anchor, TH-Z835 was retained only as a weaker comparator under explicit caution, and no third ligand was forced into the slice in the absence of adequate support. No explicit negative comparator was available. The resulting slice therefore supported a bounded anchor-versus-weaker-comparator question rather than a benchmark-complete discrimination task. Under this definition, KRAS G12D was admissible only as restricted non-pooled comparator material and not as pooled-evaluable benchmark evidence. Expanded KRAS G12D evidence-tier detail is provided in \u003cstrong\u003eSupplementary Table S5\u003c/strong\u003e and illustrated schematically in \u003cstrong\u003eSupplementary Figure S2\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 4. KRAS G12D restricted anchor-plus-weaker-comparator structure and bounded interpretation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 4. KRAS G12D restricted slice structure and bounded interpretation\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 77px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSlice\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 110px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAnchor positive\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eWeaker comparator\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 149px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eThird ligand / explicit negative\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 230px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAllowed role\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 216px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBounded verdict\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 77px;\"\u003e\n \u003cp\u003eKRAS G12D\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 110px;\"\u003e\n \u003cp\u003eMRTX1133\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eTH-Z835, with caution\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 149px;\"\u003e\n \u003cp\u003eNo third ligand included; no explicit negative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 230px;\"\u003e\n \u003cp\u003eRestricted anchor-plus-weaker-comparator slice; not pooled-metric eligible\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 216px;\"\u003e\n \u003cp\u003eAdmissible only as restricted comparator material\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cstrong\u003e3.5. CK2 is informative only as a falsification slice\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCK2 was retained not as pooled benchmark evidence, but as a falsification slice that tested rejection behavior under mechanistic contradiction (Table 5). In this case, the value of the slice did not lie in providing another positive or comparator panel. Instead, its value lay in showing that a plausible non-canonical binding hypothesis could still fail under fixed criteria when competing-site or occupancy-conflict logic was taken seriously. CK2 therefore contributed failure-mode coverage rather than benchmark breadth. Under this interpretation, the slice is scientifically useful because it shows that the framework does not only accept plausible positives; it also formalizes when a structurally attractive story should be rejected. Expanded falsification logic for the CK2 slice is provided in \u003cstrong\u003eSupplementary Table S6\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 5. CK2 falsification structure and bounded interpretation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 5. CK2 falsification-slice structure and bounded interpretation\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 77px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSlice\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 202px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eApparent hypothesis\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 211px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMechanistic conflict\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 163px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAllowed role\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePooled-metric eligible\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 230px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eVerdict\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 77px;\"\u003e\n \u003cp\u003eCK2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 202px;\"\u003e\n \u003cp\u003ePlausible non-canonical or allosteric-like binding story\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 211px;\"\u003e\n \u003cp\u003eATP-site occupancy-conflict / competing-site contradiction\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 163px;\"\u003e\n \u003cp\u003eFalsification slice only\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 230px;\"\u003e\n \u003cp\u003eInformative only as rejection and failure-mode material\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cstrong\u003e3.6. The current panel blocks pooled benchmark claims by design\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUnder the current benchmark rules, the panel does not justify pooled scoring (Figure 1). This is not a missing downstream analysis step, but the direct consequence of slice-level role assignment, row-level evaluability restriction, and exclusion of non-comparable material. PTP1B remained non-row-ready because ligand identity and applied row mapping were still constrained at the checkpoint level, KRAS G12D remained non-row-ready under strengthened strict-live-artifact requirements, TP53 Y220C was restricted to calibration-only use, CK2 was retained only as falsification material, and HIF-2\u0026alpha; remained policy-closed and outside holdout and pooled metrics. Taken together, these decisions show that the current Gen2 panel supports governed benchmark construction, but not a reviewer-defensible pooled benchmark claim.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFigure 1. Program-level pooled-scoring gate for the current Gen2 panel.\u003c/strong\u003e The current Gen2 panel contains multiple slice roles, but none meets the conditions required for pooled benchmark use at the preserved 2026-04-04 checkpoint. PTP1B and KRAS G12D remain non-row-ready, TP53 Y220C is calibration-only, CK2 is falsification material, and HIF-2\u0026alpha; is policy-closed. The resulting program state therefore supports governed benchmark construction but blocks pooled scoring by design.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.7. Cross-case synthesis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAcross the current panel, Gen2 does not behave as a simple accept-or-reject workflow. Instead, it separates candidate systems into bounded evidential roles according to what each slice can legitimately support (Table 6). In the present build, PTP1B provides the closest approach to a benchmark-like unit because it contains a same-series positive anchor, a weaker comparator, and an explicit negative, even though pooled promotion remains constrained at the preserved checkpoint. TP53 Y220C is admissible only as a limited positive-only calibration slice, KRAS G12D only as a restricted anchor-plus-weaker-comparator slice, and CK2 only as a falsification slice that tests rejection behavior under mechanistic contradiction. HIF-2\u0026alpha; remains policy-closed, while MEK1 and AKT1 reserve remain expansion candidates rather than active benchmark slices. The cross-case result is therefore not pooled benchmark performance, but a framework-level distinction between benchmark-like, calibration, restricted-comparator, falsification, and non-promoted material. That distinction is the main current benchmark result of Gen2. An expanded admissibility summary for the full panel is provided in \u003cstrong\u003eSupplementary Table S2\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 6. Cross-case synthesis of current Gen2 slice roles and admissible uses\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 6. Cross-case synthesis of current Gen2 slice roles and admissible uses\u003c/strong\u003e\u003c/p\u003e\n\u003cdiv align=\"center\"\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 67px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSlice\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRole\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 149px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAnchor or defining evidence\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eComparator structure\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 149px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eNegative or falsification logic\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 77px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePooled eligible\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 182px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCross-case verdict\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003ePTP1B\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eClosest current benchmark-like slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eFRJ anchor positive; same-series allosteric support\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eBB3 as weaker same-series comparator\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eCompound 4 as explicit same-series negative\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eClosest current approximation to a minimal benchmark unit, but not pooled-promoted at the preserved checkpoint\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eTP53 Y220C\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eLimited calibration slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003ePhiKan083 as anchor positive\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003ePK7088 as weaker comparator with caution\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo explicit negative; MB710 excluded\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eAdmissible only as positive-only calibration material\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eKRAS G12D\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eRestricted comparator slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eMRTX1133 as strong anchor\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eTH-Z835 as weaker comparator with caution\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo explicit negative; no third ligand included\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eAdmissible only as restricted anchor-plus-weaker-comparator material\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eCK2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eFalsification slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eATP-site occupancy-conflict rejection logic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNot used as a graded positive-comparator slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eMechanistic contradiction / competing-site conflict\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eInformative only as a rejection and failure-mode case\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eHIF-2\u0026alpha;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003ePolicy-closed mature-slice buildout\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eMature cavity rationale with locked anchor logic\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eComparator logic exists but not active in pooled panel\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eExcluded by policy closure\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eOutside holdout and pooled metrics at the current checkpoint\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eMEK1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eCandidate mature expansion slice\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eClean type III allosteric versus ATP-site benchmark potential\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eComparator logic externally defined\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNot promoted into active panel\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eExpansion candidate only\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eAKT1 reserve\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eReserve expansion candidate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eReserve allosteric benchmark potential\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eBaseline and comparator logic not yet promoted\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNot promoted into active panel\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eNo\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 139px;\"\u003e\n \u003cp\u003eReserve candidate only\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cem\u003eCross-case verdicts are role-based interpretations at the preserved 2026-04-04 checkpoint and do not imply pooled benchmark eligibility.\u003c/em\u003e\u003c/p\u003e"},{"header":"4. Discussion","content":"\u003cp\u003e\u003cstrong\u003e4.1. What Gen2 demonstrates at the current checkpoint\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo our knowledge, Gen2 is an early benchmark-governance framework for cryptic-pocket triage that formalizes slice-level admissibility, restricted-use roles, and pooled-scoring exclusion rules rather than treating all curated cases as uniformly benchmarkable. At the current checkpoint, Gen2 does not demonstrate pooled benchmark performance. What it demonstrates instead is something narrower but methodologically important: benchmark construction in cryptic-pocket discovery can be governed by explicit evidential rules rather than by retrospective inclusion of whatever cases appear most attractive. Under this framework, slices are not treated as automatically equivalent simply because they involve related targets or plausible ligand-pocket stories. They are separated according to what the available evidence can actually support. In the present panel, that meant retaining one slice as the closest current approximation to a minimal benchmark-like unit, restricting two others to bounded non-pooled roles, retaining one as falsification material, and keeping additional candidates outside the active panel. The main result is therefore not a benchmark score, but a disciplined demonstration that Gen2 can prevent weakly comparable, incompletely defined, or mechanistically contradictory material from being counted as benchmark evidence before reviewer-defensible conditions are met.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.2. Why blocked pooled scoring is a strength, not a failure\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA weak benchmark can be more damaging than no pooled benchmark at all. In cryptic-pocket discovery, the temptation to combine heterogeneous cases into a single apparent performance summary is strong, particularly when some systems carry plausible ligand-pocket stories, some have partial structural support, and some have only restricted comparator logic. The present study took the opposite approach. Instead of forcing every curated slice into pooled evaluation, Gen2 treated non-comparability itself as an actionable result. Under that rule, blocked pooled scoring is not evidence of framework failure. It is evidence that the admission criteria are functioning as intended.\u003c/p\u003e\n\u003cp\u003eThis distinction matters because a pooled benchmark number is only meaningful if the underlying rows are genuinely comparable. If non-row-ready slices, calibration-only material, policy-closed systems, and falsification cases are allowed to enter the same pooled analysis, the resulting score may look quantitative while resting on a mixed evidential base. Such a result would be more vulnerable to reviewer criticism than a deliberately blocked benchmark, because it would confuse administrative inclusion with evaluable evidence. By keeping those categories separate, the current framework preserves a clear boundary between what can support a benchmark claim and what can only support restricted interpretation.\u003c/p\u003e\n\u003cp\u003eThe current checkpoint therefore illustrates a conservative but scientifically preferable outcome. The present checkpoint should be interpreted as insufficient available evidence for pooled benchmarking, rather than as failure of the Gen2 framework. Gen2 does not reward apparent panel breadth at the expense of evidential discipline. Instead, it treats refusal to over-score the panel as part of benchmark quality control. In practical terms, this means that the present manuscript contributes a benchmark framework that is willing to stop before pooled evaluation when the slice set does not justify it. That restraint is methodologically stronger than reporting an attractive benchmark number built from cases that were not ready to be compared on equal terms.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.3. What the current panel still lacks and how it can be strengthened\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe present panel is structured enough to support a benchmark-construction paper, but it is not yet broad enough to support a mature pooled benchmark claim. The main limitation is not simply the number of targets. It is the incomplete coverage of benchmark functions. At the current checkpoint, the panel contains one slice that most closely approaches a benchmark-like triage unit, two slices that are restricted to bounded non-pooled roles, one falsification slice, and no mature pooled-evaluable set. This means that the program still lacks a sufficiently diverse collection of row-ready slices with clean comparator structure, explicit admissibility logic, and comparable evidential maturity.\u003c/p\u003e\n\u003cp\u003eThe most useful next strengthening step is therefore not to add more targets indiscriminately, but to add slices that fill missing benchmark functions. In this context, mature buildout candidates such as HIF-2α and MEK1 are valuable because they address gaps that are not covered by the current panel. HIF-2α offers a non-kinase internal-cavity case with stronger mature-slice potential, while MEK1 offers a cleaner type III versus ATP-site discrimination problem. By contrast, reserve candidates should remain outside the active panel until their baseline state, comparator logic, and slice definitions are fixed to the same standard. Under this logic, benchmark expansion should be driven by missing evidential function rather than by target novelty alone. Candidate mature-slice expansion logic is summarized in \u003cstrong\u003eSupplementary Table S7\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eA second strengthening route is internal rather than expansive. The current panel would also become materially stronger if restricted slices were upgraded to row-ready status under the same fixed rules already used here. In practice, that means resolving the remaining constraints that block pooled use, rather than relaxing the admission threshold. This is especially important because the current paper is stronger when it shows that the same evidential standards apply both to slice admission and to later attempts at panel expansion. The benchmark therefore does not need arbitrary growth. It needs either additional mature slices that fill missing benchmark roles, or stricter maturation of the slices already present.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.4. Limitations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSeveral limitations define the evidential ceiling of the present study. First, the current panel does not support pooled benchmark claims. That restriction is intentional, but it also means that the manuscript cannot yet provide a cross-slice performance summary, comparative accuracy statistic, or broad generality claim for Gen2. The strongest contribution at this stage is therefore framework-level governance rather than completed benchmark performance. This should be interpreted as insufficient available evidence for pooled benchmarking, rather than as failure of the framework itself.\u003c/p\u003e\n\u003cp\u003eSecond, the panel remains asymmetric in slice maturity and evidential structure. PTP1B comes closest to a benchmark-like unit because it contains a same-series anchor positive, a weaker comparator, and an explicit negative, but even that slice remains constrained at the preserved checkpoint. TP53 Y220C and KRAS G12D were retained only in restricted non-pooled roles, while CK2 was informative only as falsification material. This asymmetry is scientifically manageable, but it limits how far cross-case comparison can be pushed.\u003c/p\u003e\n\u003cp\u003eThird, not all restricted slices fail for the same reason. Some remain limited because of missing explicit negatives, some because of row-readiness constraints, some because of unresolved mapping or comparator structure, and some because their value lies mainly in rejection behavior rather than benchmark inclusion. As a result, the current study should not be read as showing that all non-pooled slices are equally immature or equally informative. Their restricted status reflects different evidential bottlenecks.\u003c/p\u003e\n\u003cp\u003eFourth, the benchmark remains dependent on literature-supported and structure-supported slice definitions rather than on a fully matured holdout panel. This is appropriate for a benchmark-construction paper, but it means that some conclusions are still conditional on future slice maturation under the same fixed rules. In particular, the present study does not yet show how Gen2 behaves across a larger set of row-ready slices spanning multiple target classes under a common pooled scoring framework.\u003c/p\u003e\n\u003cp\u003eFinally, the current manuscript does not integrate blinded orthogonal reanalysis or experimental assay follow-up into the benchmark core. That boundary was deliberate, because mixing benchmark construction with validation and assay layers at this stage would have weakened interpretability. However, it also means that the present paper addresses benchmark admissibility rather than full end-to-end validation. Future work will therefore need to test whether the same governance logic remains robust when additional mature slices are promoted and when benchmark-ready slices are challenged under external validation conditions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4.5. Conclusion\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study presents Gen2 not as a completed pooled benchmark, but as a governed benchmark-construction framework for binding hypothesis triage in cryptic-pocket discovery. The main result is that benchmark assembly itself can be treated as a scientific decision layer, with explicit rules for slice admission, role assignment, row registration, pooled-metric eligibility, and falsification handling. Under those rules, the current panel did not justify pooled scoring. Rather than weakening the framework, that outcome showed that Gen2 can prevent weakly comparable, incompletely defined, or mechanistically contradictory cases from being counted as benchmark evidence before reviewer-defensible conditions are met.\u003c/p\u003e\n\u003cp\u003eAt the present checkpoint, Gen2 supports one slice that most closely approaches a minimal benchmark-like unit, two slices retained only in restricted non-pooled roles, one falsification slice, and additional mature-slice candidates that remain outside the active panel. Taken together, these assignments define the current contribution of the framework: not a benchmark number, but a reproducible evidential architecture for deciding what may and may not count in a cryptic-pocket benchmark. That is the central conclusion of the present work.\u003c/p\u003e\n\u003cp\u003eThe next stage is clear. Gen2 will become a stronger benchmark only if additional slices are matured under the same fixed rules, or if currently restricted slices become row-ready without relaxing admission standards. Until then, the appropriate claim is bounded but meaningful: Gen2 provides a reviewer-defensible framework for benchmark construction in cryptic-pocket discovery, and its present value lies in governed admissibility rather than in pooled performance.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003e5. Acknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6. Funding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no external funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e7. Author Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHoosdally Shakeel conceived the study, designed the benchmark framework, curated the benchmark slices, performed the analysis, interpreted the results, prepared the figures and tables, and wrote the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e8. Conflicts of Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author declares a patent-related competing interest. Intellectual property related to aspects of the framework described in this manuscript has been filed through the USPTO.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e9. Data Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data supporting the findings of this study are available from the author on reasonable request. These materials include the benchmark tables, slice-role assignments, row-registration outputs, figure source files, and supporting benchmark-construction records needed to interpret the results reported in this manuscript. Because aspects of the framework are the subject of patent-related protection, full release of all underlying workflow materials, scripts, and intermediate development records is not included in the present manuscript. Any shared materials will therefore be limited to those necessary to support transparency, interpretation, and reproducibility of the reported benchmark state without disclosing protected implementation details.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eVajda S, Beglov D, Wakefield AE, Egbert M, Whitty A (2018) Cryptic binding sites on proteins: definition, detection, and druggability. Curr Opin Chem Biol 44:1\u0026ndash;8. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.cbpa.2018.05.003\u003c/span\u003e\u003cspan address=\"10.1016/j.cbpa.2018.05.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eŠkrh\u0026aacute;k V, Novotn\u0026yacute; M, Feidakis CP, Kriv\u0026aacute;k R, Hoksza D (2025) CryptoBench: cryptic protein-ligand binding sites dataset and benchmark. Bioinformatics 41(1):btae745. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bioinformatics/btae745\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/btae745\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShakeel H (2026) Gen2: A Biophysical Triage Framework for Binding Hypotheses in Cryptic Pocket Discovery. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.26434/chemrxiv.15001255/v1\u003c/span\u003e\u003cspan address=\"10.26434/chemrxiv.15001255/v1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. ChemRxiv\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBurley SK, Bhikadiya C, Bi C, Bittrich S, Chen L, Crichlow GV et al (2023) RCSB Protein Data Bank (RCSB.org): delivery of experimentally determined PDB structures alongside one million computed structure models of proteins from artificial intelligence/machine learning. Nucleic Acids Res 51(D1):D488\u0026ndash;D508. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/nar/gkac1077\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkac1077\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSayers EW, Beck J, Bolton EE, Bourexis D, Brister JR, Canese K et al (2025) Database resources of the National Center for Biotechnology Information. Nucleic Acids Res 53(D1):D20\u0026ndash;D32. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/nar/gkae988\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkae988\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKooistra AJ, Kanev GK, van Linden OPGJ, Leurs R, de Esch IJP, de Graaf C (2016) KLIFS: a structural kinase-ligand interaction database. Nucleic Acids Res 44(D1):D365\u0026ndash;D371. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/nar/gkv1082\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkv1082\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe J, Liu X, Zhu C, Zha J, Ou W, Zhang G et al (2024) ASD2023: towards the integrating landscapes of allosteric knowledgebase. Nucleic Acids Res 52(D1):D376\u0026ndash;D383. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/nar/gkad915\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkad915\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWiesmann C, Barr KJ, Kung J, Zhu J, Erlanson DA, Shen W et al (2004) Allosteric inhibition of protein tyrosine phosphatase 1B. Nat Struct Mol Biol 11(8):730\u0026ndash;737. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nsmb803\u003c/span\u003e\u003cspan address=\"10.1038/nsmb803\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKrishnan N, Koveal D, Miller DH, Xue B, Akshinthala SD, Kragelj J et al (2014) Targeting the disordered C terminus of PTP1B with an allosteric inhibitor. Nat Chem Biol 10(7):558\u0026ndash;566. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nchembio.1528\u003c/span\u003e\u003cspan address=\"10.1038/nchembio.1528\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoeckler FM, Joerger AC, Jaggi G, Rutherford TJ, Veprintsev DB, Fersht AR (2008) Targeted rescue of a destabilized mutant of p53 by an in silico screened drug. Proc Natl Acad Sci U S A 105(30):10360\u0026ndash;10365. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1073/pnas.0805326105\u003c/span\u003e\u003cspan address=\"10.1073/pnas.0805326105\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang X, Allen S, Blake JF, Bowcut V, Briere DM, Calinisan A et al (2022) Identification of MRTX1133, a noncovalent, potent, and selective KRASG12D inhibitor. J Med Chem 65(4):3123\u0026ndash;3133. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1021/acs.jmedchem.1c01688\u003c/span\u003e\u003cspan address=\"10.1021/acs.jmedchem.1c01688\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMao Z, Xiao H, Shen P, Yang Y, Guan Y, Wang Y et al (2022) KRAS(G12D) can be targeted by potent inhibitors via formation of a salt bridge. Cell Discov 8:5. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41421-021-00368-w\u003c/span\u003e\u003cspan address=\"10.1038/s41421-021-00368-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWallace EM, Rizzi JP, Han G, Wehn PM, Cao Z, Du X et al (2016) A small-molecule antagonist of HIF2α is efficacious in preclinical models of renal cell carcinoma. Cancer Res 76(18):5491\u0026ndash;5500. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1158/0008-5472.CAN-16-0473\u003c/span\u003e\u003cspan address=\"10.1158/0008-5472.CAN-16-0473\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRen X, Wang J, Cui S, Gao B, Li M, Sun Y et al (2022) Structural basis for the allosteric inhibition of hypoxia-inducible factor HIF-2 by belzutifan. Mol Pharmacol 102(6):366\u0026ndash;376. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1124/molpharm.122.000525\u003c/span\u003e\u003cspan address=\"10.1124/molpharm.122.000525\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHatzivassiliou G, Haling JR, Chen H, Song K, Price S, Heald R et al (2013) Mechanism of MEK inhibition determines efficacy in mutant KRAS- versus BRAF-driven cancers. Nature 501(7466):232\u0026ndash;236. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nature12441\u003c/span\u003e\u003cspan address=\"10.1038/nature12441\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakano K, Munehira Y, Itou J, Watanabe N, Horiguchi M, Takikawa S et al (2023) Discovery of a novel ATP-competitive MEK inhibitor DS03090629 that overcomes resistance conferred by BRAF overexpression in BRAF-mutated melanoma. Mol Cancer Ther 22(3):317\u0026ndash;329. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1158/1535-7163.MCT-22-0306\u003c/span\u003e\u003cspan address=\"10.1158/1535-7163.MCT-22-0306\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"cryptic pocket, allostery, benchmark construction, benchmark governance, binding hypothesis triage, slice admissibility, pooled scoring, falsification slice, PTP1B, TP53 Y220C, KRAS G12D","lastPublishedDoi":"10.21203/rs.3.rs-9336311/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9336311/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBenchmarking in cryptic-pocket and allosteric discovery is often weakened by forcing heterogeneous case studies into pooled scoring despite ambiguous labels, unstable site assignment, missing row-level outputs, or mismatched evidential standards. Here, we present Gen2 as a governance-first benchmark framework for binding hypothesis triage in cryptic pocket discovery. Rather than treating benchmark assembly as a secondary administrative step, Gen2 treats it as part of the scientific method: each candidate slice is screened against frozen evidential rules, assigned a bounded role, and either admitted, parked, excluded, or retained as calibration or falsification material before pooled evaluation is considered. Applying this framework to the current panel produced a preserved no-active-slice-open checkpoint. Under these rules, HIF-2α remained policy-closed, TP53 Y220C remained calibration-only, CK2 was retained as falsification material, and KRAS G12D and PTP1B remained non-row-ready for different reasons. The principal result is therefore not pooled benchmark performance, but demonstration that Gen2 prevents invalid pooled claims by blocking premature scoring and preserving only reviewer-defensible evaluable units. This establishes a reproducible benchmark-construction layer for future multi-slice evaluation once row-ready systems and explicit row mappings are available.\u003c/p\u003e","manuscriptTitle":"Gen2: Building a Reviewer-Defensible Benchmark for Binding Hypothesis Triage in Cryptic Pocket Discovery","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-09 16:03:07","doi":"10.21203/rs.3.rs-9336311/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2b2a2334-4646-4539-8043-e5062e240bd3","owner":[],"postedDate":"April 9th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":65803732,"name":"Computational Biology"}],"tags":[],"updatedAt":"2026-04-09T16:03:07+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-09 16:03:07","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9336311","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9336311","identity":"rs-9336311","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.