<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">79663</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.079663</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>CF2-SLAM: Conformal-Calibrated Foundation-Factor Graph SLAM across Modalities and Domains</article-title>
<alt-title alt-title-type="left-running-head">CF2-SLAM: Conformal-Calibrated Foundation-Factor Graph SLAM across Modalities and Domains</alt-title>
<alt-title alt-title-type="right-running-head">CF2-SLAM: Conformal-Calibrated Foundation-Factor Graph SLAM across Modalities and Domains</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Chen</surname><given-names>Xiangqin</given-names></name><email>rexc@alumni.psu.edu</email></contrib>
<aff id="aff-1"><institution>College of Engineering, Pennsylvania State University</institution>, <addr-line>University Park, PA</addr-line>, <country>USA</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Xiangqin Chen. Email: <email>rexc@alumni.psu.edu</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>39</elocation-id>
<history>
<date date-type="received">
<day>26</day>
<month>01</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Author. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Author</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_79663.pdf"></self-uri>
<abstract>
<p>Simultaneous localization and mapping (SLAM) must remain reliable when sensing suites and operating conditions vary across platforms and deployments. Beyond correspondence degradation, a dominant deployment failure mode is <italic>misweighted</italic> constraints: under distribution shift, uncertainty estimates can become miscalibrated, allowing a small set of overconfident factors to dominate iterative optimization and destabilize inference. This article presents conformal-calibrated foundation-factor graph SLAM (<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula>), a sensor-agnostic framework that combines frozen foundation representations with lightweight probabilistic factor heads that emit explicit residuals and covariances, and a classical factor-graph back-end for principled multi-modal fusion. To mitigate systematic misweighting under shift, an online conformal calibration layer is introduced to rescale factor covariances by aligning empirical residual quantiles with target quantiles on a per-factor-family basis. Loop closure is further integrated through foundation-descriptor retrieval for candidate proposal and conservative geometric verification for graph insertion, controlling false loop constraints without relying on dataset-specific place-recognition supervision. Across heterogeneous benchmarks spanning monocular, stereo, red-green-blue-depth (RGB-D), and visual-inertial settings, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula> operates without retraining and shows improved robustness trends under zero-shot transfer, consistent with stabilized factor weighting.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Simultaneous localization and mapping (SLAM)</kwd>
<kwd>factor graph optimization</kwd>
<kwd>foundation models</kwd>
<kwd>uncertainty estimation</kwd>
<kwd>online calibration</kwd>
<kwd>conformal calibration</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Simultaneous localization and mapping (SLAM) is a core state-estimation module for embodied platforms ranging from aerial robots to autonomous vehicles and augmented reality/virtual reality (AR/VR) devices. In practice, deployments vary in sensing configuration (monocular/stereo/red-green-blue-depth (RGB-D)/inertial measurement unit (IMU)) and operating conditions (illumination, weather, scene layout, motion, and dynamics), inducing distribution shifts that degrade robustness.</p>
<p>Classical geometric pipelines remain attractive due to transparent objectives and auditable components, but they rely on brittle data association and typically assume stationary noise models. Learned SLAM mitigates perceptual brittleness via learned representations and stronger visual priors [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>], yet a less visible failure mode often dominates in deployment: <italic>miscalibrated confidence under shift</italic>. In iterative back-ends (bundle adjustment or factor graphs), relative factor weighting controls conditioning and convergence. A small subset of overconfident, incorrect constraints can dominate the normal equations and cause divergence or persistent bias, especially when mixing factor types of different dimensions and noise profiles.</p>
<p>The target of this work is a unified SLAM framework that operates across modalities while retaining the interpretability of a probabilistic back-end. Two principles guide the design. First, frozen foundation models can provide more transferable representations than task-specific encoders [<xref ref-type="bibr" rid="ref-3">3</xref>]. Second, if learned components emit explicit residuals and covariances, inference can be posed as a classical factor-graph maximum a posteriori (MAP) problem [<xref ref-type="bibr" rid="ref-4">4</xref>], enabling principled fusion. However, transferability alone does not prevent systematic misweighting under shift. An online conformal-style calibration mechanism is therefore introduced to adjust factor covariance magnitudes using residual quantiles [<xref ref-type="bibr" rid="ref-5">5</xref>], aiming to support more reliable optimization behavior over time under distribution shift.</p>
<p>Because widely used SLAM datasets differ in available sensor fields and calibration metadata, <xref ref-type="sec" rid="s5">Section 5</xref> documents modality availability and <xref ref-type="sec" rid="s6_2">Section 6.2</xref> fixes evaluation protocols to reduce inadvertent modality leakage.</p>
<p><italic>Contributions</italic>.</p>
<p>This article makes three contributions: (1) <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula> is formulated as a factor-graph SLAM framework in which frozen foundation features feed lightweight probabilistic factor heads, producing explicit residuals and covariances in a transparent MAP objective; (2) an online conformal calibration layer rescales factor covariances by matching observed residual quantiles to target quantiles, mitigating systematic misweighting under domain shift during sequential optimization; (3) descriptor-topological loop closure, i.e., descriptor-space candidate proposal followed by geometric verification and loop-closure (LC)-factor insertion, uses foundation descriptors for candidate proposal and geometric verification for conservative graph insertion, avoiding dataset-specific place-recognition supervision while controlling false loop closures.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Classical geometric SLAM and visual-inertial odometry (VIO) remain strong baselines when sensing assumptions are matched to deployment. Feature-based systems such as ORB-SLAM3 [<xref ref-type="bibr" rid="ref-6">6</xref>] remain highly competitive, while recent geometry-aware learned systems such as Photo-SLAM [<xref ref-type="bibr" rid="ref-1">1</xref>] and IBD-SLAM [<xref ref-type="bibr" rid="ref-2">2</xref>] illustrate the continued value of combining stronger visual representations with explicit optimization back-ends. However, their accuracy and stability still depend strongly on data association quality and on appropriately tuned noise models across visual and inertial factors.</p>
<p>Learned front-ends improve correspondence robustness under viewpoint, illumination, and texture variation. Representative recent examples include DINOv2 [<xref ref-type="bibr" rid="ref-3">3</xref>], LightGlue [<xref ref-type="bibr" rid="ref-7">7</xref>], RoMa [<xref ref-type="bibr" rid="ref-8">8</xref>], and IBD-SLAM [<xref ref-type="bibr" rid="ref-2">2</xref>]. These methods show that learned representations and learned residual models can substantially strengthen SLAM front-ends, but they also expose a recurring weakness: confidence or covariance estimates that are reliable in-domain can become miscalibrated under cross-dataset, cross-sensor, or synthetic-to-real transfer.</p>
<p>Loop closure is commonly structured as candidate retrieval followed by geometric validation before graph insertion. In this setting, the proposed descriptor-topological module follows the same principle: foundation descriptors are used only for candidate proposal, while accepted loop constraints are instantiated as verified LC factors after dense geometric checking. This is distinct from metric-semantic SLAM in the sense of explicit object- or scene-level labeling, and it is also distinct from recent dense mapping systems such as Gaussian Splatting SLAM [<xref ref-type="bibr" rid="ref-9">9</xref>] and SplaTAM [<xref ref-type="bibr" rid="ref-10">10</xref>], which emphasize reconstruction quality and often assume RGB-D inputs and higher compute budgets.</p>
<p>Recent work on post-hoc neural calibration and conformal prediction has strengthened uncertainty quantification in supervised learning [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>], but their role in SLAM is less straightforward because SLAM involves sequential correlation, heterogeneous factor dimensions, and solver-level sensitivity to relative factor weighting. The emphasis here is therefore not only uncertainty reporting but solver conditioning under distribution shift. A related perspective on resilience under changing operating conditions appears in graph-structured logistics routing, where learned spatiotemporal risk prediction is integrated with dynamic edge weighting to maintain robust decision making under congestion and demand fluctuations [<xref ref-type="bibr" rid="ref-13">13</xref>]. The proposed online conformal layer complements standard robust losses: robust losses suppress large instantaneous outliers, whereas conformal rescaling corrects systematic covariance scale mismatch across factor families so that no single miscalibrated modality dominates the normal equations after transfer.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Problem Formulation</title>
<p>This work considers heterogeneous sensor streams (monocular/stereo/RGB-D/IMU) and estimates a trajectory (and optionally map parameters) over a horizon. Let the state at time <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>t</mml:mi></mml:math></inline-formula> be
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">T</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mtext>SE</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mrow><mml:mtext mathvariant="bold">v</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mn>3</mml:mn></mml:msup><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mrow><mml:mtext mathvariant="bold">b</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mn>6</mml:mn></mml:msup><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">T</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> is the pose, <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">v</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> velocity, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">b</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> IMU biases (if applicable), and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> optional map parameters.</p>
<p>A factor graph over <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> defines the MAP objective:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow></mml:mrow></mml:munder><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where factor <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>k</mml:mi></mml:math></inline-formula> contributes residual <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, covariance matrix <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>, and robust loss <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-4">4</xref>]. The standard convention is followed that <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is a covariance matrix; correspondingly, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is the information (weight) matrix used inside <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>. Beyond uncertainty reporting, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> directly controls factor influence and conditioning; thus systematic miscalibration under shift can destabilize optimization. In <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula>, each learned factor family predicts an initial <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>, and the online conformal calibration layer rescales <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> (hence <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>) to improve weighting reliability across modalities and domains. The method below retains interpretability of <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref> while improving weighting reliability across modalities and domains. Throughout the paper, foundation-feature (FF), depth/disparity (D), inertial measurement unit (IMU), and loop-closure (LC) denote the four factor families used in the MAP objective. FF/D/LC correspond to visual or verified loop constraints whose residual blocks are paired with learned or learned-assisted covariance estimates, while IMU denotes the standard analytic preintegration factor with its usual inertial covariance when IMU measurements are available.</p>
</sec>
<sec id="s4">
<label>4</label>
<title>Method: <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula></title>
<sec id="s4_1">
<label>4.1</label>
<title>Overview</title>
<p><inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula> pairs a learned front-end with a classical factor-graph back-end. A frozen foundation encoder produces transferable representations; lightweight factor heads emit probabilistic constraints (residuals and covariances); an online calibration layer rescales covariances per factor type; inference uses Gauss-Newton/Levenberg-Marquardt (GN/LM) on a sliding-window graph. Modality changes do not require retraining: analytic factors (e.g., IMU preintegration) are enabled only when the corresponding sensor fields exist. An overview is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>System overview of conformal-calibrated foundation-factor graph SLAM (<inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula>). Heterogeneous sensors <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> frozen foundation features <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> probabilistic factor heads <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> online conformal calibration <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> factor-graph optimization for trajectory and optional map.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-1.tif"/>
</fig>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Foundation Representations</title>
<p>A frozen foundation model <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> (e.g., DINOv2 [<xref ref-type="bibr" rid="ref-3">3</xref>]) is used to compute (i) dense token features <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">F</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and (ii) a global descriptor <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mi>C</mml:mi></mml:msup></mml:math></inline-formula> for retrieval. Token-wise normalization improves stability for similarity-based residuals. Freezing <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>&#x03D5;</mml:mi></mml:math></inline-formula> limits dataset-specific overfitting and focuses trainable capacity on small heads that translate features into residual and covariance predictions.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Graph Construction and Edge Selection</title>
<p>Optimization is performed over a sliding window of <italic>N</italic> states with optional loop-closure edges. Temporal edges ensure local observability; sparse covisibility edges add redundancy without quadratic connectivity. Loop candidates are proposed by retrieving neighbors in descriptor space and inserted only after geometric verification (<xref ref-type="sec" rid="s4_5">Section 4.5</xref>), since false high-confidence loops can bias the entire graph.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Probabilistic Learned Factors</title>
<p>Factor types and modality dependencies are summarized in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>; only factors supported by dataset fields (<xref ref-type="sec" rid="s5">Section 5</xref>) are enabled to avoid modality leakage.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Unified factor graph for <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula>. Edge availability depends on sensor modalities (monocular/stereo/red-green-blue-depth (RGB-D)/inertial measurement unit (IMU)) and verified loop closures.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-2.tif"/>
</fig>
<p>Four factor families are used: a foundation-feature (FF) factor, a depth/disparity (D) factor when depth is present, an IMU preintegration factor when IMU exists [<xref ref-type="bibr" rid="ref-14">14</xref>], and a loop-closure (LC) factor after verification. The descriptor-topological loop-closure module described below is therefore not a fifth factor family; rather, it is the proposal-and-verification pipeline whose accepted outputs are instantiated as LC factors in the graph. For FF, given an estimated relative pose <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">T</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and optional depth <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">D</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, pixel <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>u</mml:mi></mml:math></inline-formula> is warped from <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>i</mml:mi></mml:math></inline-formula> to <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>j</mml:mi></mml:math></inline-formula> and features are compared:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>FF</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>u</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">F</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>u</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">F</mml:mtext></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">T</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:mi mathvariant="normal">&#x03A0;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">D</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>For RGB-D depth consistency:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msubsup><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>D</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>u</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>u</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>D</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo stretchy="false">&#x2190;</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>u</mml:mi><mml:mo>;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">T</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>and for stereo an analogous disparity form applies. For loop closure, after verification a relative pose constraint is added in <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mrow><mml:mi mathvariant="fraktur">s</mml:mi><mml:mi mathvariant="fraktur">e</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>LC</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">T</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">T</mml:mtext></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mtext mathvariant="bold">T</mml:mtext></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mn>6</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The IMU family uses the standard preintegrated residual over pose, velocity, and bias states; <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:msub></mml:math></inline-formula> in the calibration section below always denotes the dimension of the corresponding residual block for factor family <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>.</p>
<p>Each factor head outputs a positive-definite covariance <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>. Diagonal <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>diag</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">s</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> (log-variances <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">s</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula>) or a Cholesky parameterization is used when needed. Robust losses <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>&#x03C1;</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> (e.g., Huber/Cauchy) reduce sensitivity to sporadic outliers; calibration (below) targets systematic misweighting under shift. In particular, the covariance parameterization is chosen so that the initial <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is symmetric positive definite (SPD); the robust loss is applied to the Mahalanobis energy in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>, whereas the conformal layer only rescales the covariance magnitude for a factor family and does not change the residual definition itself.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Descriptor-Topological Loop Closure with Geometric Verification</title>
<p>Loop closure separates candidate proposal from constraint insertion (<xref ref-type="fig" rid="fig-3">Fig. 3</xref>). In this article, &#x201C;descriptor-topological&#x201D; refers to this retrieval-and-verification pipeline, while the actual graph element added after a successful check is the LC factor defined above. Candidates are retrieved using <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and a memory bank; verification uses dense matching (e.g., LoFTR [<xref ref-type="bibr" rid="ref-15">15</xref>]) and robust Perspective-n-Point (PnP)/SE(3) estimation. Only verified loops are inserted, controlling false-loop contamination while avoiding dataset-specific place-recognition supervision.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Descriptor-topological loop closure. Foundation descriptors propose candidates; dense matching verifies geometry; accepted closures become loop-closure (LC) factors in the graph.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-3.tif"/>
</fig>
<p>To balance accuracy and efficiency, proposal and verification are decoupled: descriptor retrieval generates hypotheses, and dense matching with robust pose estimation acts as a conservative gate, executed sparsely (e.g., on keyframes) with a capped number of candidates per query; the reported frame rate, measured in frames per second (FPS), includes the amortized verification cost. Descriptor retrieval alone is high-recall but insufficiently conservative for graph insertion under perceptual aliasing and repeated structures; dense geometric verification is therefore required before adding a loop-closure factor.</p>
<p>In dynamic scenes, non-stationarity mainly impacts retrieval and correspondence; the verification gate rejects inconsistent hypotheses before graph insertion, and dynamic-aware masking or temporal-consistency checks can be incorporated when needed.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Online Conformal Calibration of Factor Covariances</title>
<p>Under domain shift, predicted uncertainties can become systematically miscalibrated, altering factor influence in GN/LM. Covariances are therefore rescaled online using residual statistics (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>).</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Online conformal calibration. For each factor type, match observed residual quantiles to target quantiles and rescale covariances accordingly.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-4.tif"/>
</fig>
<p>For factor <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>k</mml:mi></mml:math></inline-formula>, define the score
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>s</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:msqrt><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>For each factor type <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>, a window of scores is maintained and the empirical <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> quantile <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msubsup><mml:mi>q</mml:mi><mml:mrow><mml:mi>&#x03C4;</mml:mi></mml:mrow><mml:mrow><mml:mtext>obs</mml:mtext></mml:mrow></mml:msubsup></mml:math></inline-formula> is computed. If well calibrated and approximately Gaussian, then <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msubsup><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:math></inline-formula> follows <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msup><mml:mi>&#x03C7;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> with residual dimension <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:msub></mml:math></inline-formula>. The target quantile is set to <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msubsup><mml:mi>q</mml:mi><mml:mrow><mml:mi>&#x03C4;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>tar</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msqrt><mml:msubsup><mml:mi>&#x03C7;</mml:mi><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:msub></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:msqrt></mml:math></inline-formula> and the covariance is rescaled as:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">&#x2190;</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:msub><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msubsup><mml:mi>q</mml:mi><mml:mrow><mml:mi>&#x03C4;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>obs</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>q</mml:mi><mml:mrow><mml:mi>&#x03C4;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>tar</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>This increases covariance when residuals are larger-than-expected and decreases it when residuals are smaller-than-expected. Warm-up and caps on <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:msub></mml:math></inline-formula> are used to reduce sensitivity to early transients and abrupt regime changes. Because <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mi>&#x03C4;</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, the rescaled covariance remains SPD whenever the initial <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> is SPD; the update therefore changes only the overall scale of a factor family, not its internal correlation structure. In this implementation, robust losses and conformal rescaling play complementary roles: the robust loss suppresses large instantaneous outliers at the factor level, whereas conformal rescaling corrects slower covariance-scale mismatch accumulated over a window. It is also emphasized that <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref> is used here as an online conformal-style calibration rule rather than as a strict conformal-prediction guarantee. Sequential SLAM violates exchangeability through temporal correlation, state-feedback, and changing operating regimes, so the <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msup><mml:mi>&#x03C7;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula> target should be interpreted as an approximate reference for residual-scale matching rather than a finite-sample coverage guarantee.</p>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Training Details</title>
<p>Factor heads are trained in two stages, with pretraining on TartanAir to initialize residual and covariance behavior. Optimization uses AdamW (learning rate <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, batch size 16, 100k iterations). A probabilistic objective aligned with inference is minimized:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>k</mml:mi></mml:munder><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mtext mathvariant="bold">r</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>aux</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>aux</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> includes smoothness priors and self-supervised consistency terms. All experiments use NVIDIA RTX 4090 graphics processing units (GPUs).</p>
<p>The inference procedure is summarized in Algorithm 1.</p>
<fig id="fig-10">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-10.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Datasets and Protocols</title>
<p>Sensor fields are documented to support cross-modal comparisons, and only factors supported by available measurements are enabled. Evaluation is conducted on seven public benchmarks that jointly cover outdoor driving (KITTI Odometry, KITTI-360), aerial visual-inertial simultaneous localization and mapping on EuRoC Micro Aerial Vehicle (MAV), and indoor RGB-D tracking/relocalization/mapping (TUM RGB-D, ScanNet, 7-Scenes), with TartanAir used for broad pretraining and stress-testing under diverse simulated conditions. Specifically, KITTI/KITTI-360 provide rectified stereo driving sequences for odometry and loop closure under appearance change; EuRoC provides synchronized stereo &#x002B; IMU for visual-inertial odometry (VIO) evaluation under aggressive motion; TUM RGB-D and 7-Scenes emphasize indoor tracking and relocalization with depth; ScanNet supports large-scale indoor RGB-D tracking and dense reconstruction; TartanAir supplies diverse environments and modalities to initialize factor heads for transfer. Dataset fields and typical tasks are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Dataset fields and supported tasks (based on official documentation). &#x201C;RGB&#x201D; denotes the color image stream, &#x201C;RGB-D&#x201D; denotes red-green-blue-depth, &#x201C;IMU&#x201D; denotes inertial measurement unit, &#x201C;GT Traj&#x201D; indicates ground-truth (GT) trajectory/state sufficient for absolute trajectory error (ATE)/relative pose error (RPE), &#x201C;Intr.&#x201D; indicates published intrinsics/calibration, and &#x201C;SLAM&#x201D; denotes simultaneous localization and mapping.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Dataset</th>
<th align="center">RGB</th>
<th align="center">Stereo</th>
<th align="center">Depth</th>
<th align="center">IMU</th>
<th align="center">Intr.</th>
<th align="center">GT Traj/State</th>
<th align="center">Typical SLAM Tasks</th>
</tr>
</thead>
<tbody>
<tr>
<td>KITTI Odometry [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>&#x2713; (00&#x2013;10)</td>
<td>Visual odometry/SLAM, loop closure</td>
</tr>
<tr>
<td>KITTI-360 [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>(global positioning system/inertial navigation system)</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>Long-term odometry, mapping</td>
</tr>
<tr>
<td>EuRoC MAV [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>Visual-inertial odometry (VIO), loop closure</td>
</tr>
<tr>
<td>TUM RGB-D [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>(robot operating system (ROS) bag)</td>
<td>&#x2713;</td>
<td>&#x2713; (most)</td>
<td>RGB-D SLAM, dense mapping</td>
</tr>
<tr>
<td>ScanNet [<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>&#x2713; (poses)</td>
<td>Tracking, dense reconstruction</td>
</tr>
<tr>
<td>7-Scenes</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>&#x2713;</td>
<td>&#x2013;</td>
<td>default depth intrinsics</td>
<td>&#x2713;</td>
<td>Relocalization</td>
</tr>
<tr>
<td>TartanAir</td>
<td>&#x2713;</td>
<td>(varies)</td>
<td>&#x2713;</td>
<td>(varies)</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>Pretraining, stress tests</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6">
<label>6</label>
<title>Experimental Setup</title>
<sec id="s6_1">
<label>6.1</label>
<title>Baselines</title>
<p>Baselines and modality requirements are summarized in <xref ref-type="table" rid="table-2">Table 2</xref>. Comparison is made only where required fields exist (<xref ref-type="table" rid="table-1">Table 1</xref>), and runtime/hardware is reported when available.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Baseline systems and modality requirements. Only compare on datasets where required fields exist (see <xref ref-type="table" rid="table-1">Table 1</xref>). &#x201C;SLAM&#x201D; denotes simultaneous localization and mapping, &#x201C;VO&#x201D; denotes visual odometry, &#x201C;RGB-D&#x201D; denotes red-green-blue-depth, &#x201C;IMU&#x201D; denotes inertial measurement unit, and &#x201C;mono&#x201D; denotes monocular.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Category</th>
<th>Methods</th>
<th>Required Fields</th>
</tr>
</thead>
<tbody>
<tr>
<td>Feature-based SLAM</td>
<td>ORB-SLAM3 [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>Mono/stereo/RGB-D, intrinsics; optional IMU</td>
</tr>
<tr>
<td>Direct VO/SLAM</td>
<td>DSO [<xref ref-type="bibr" rid="ref-22">22</xref>], SVO2 [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Mono/stereo, intrinsics</td>
</tr>
<tr>
<td>Visual-Inertial</td>
<td>OKVIS [<xref ref-type="bibr" rid="ref-24">24</xref>], VINS-Mono [<xref ref-type="bibr" rid="ref-14">14</xref>], VINS-Fusion</td>
<td>Stereo/mono &#x002B; IMU, intrinsics, IMU calibration</td>
</tr>
<tr>
<td>Deep VO/SLAM</td>
<td>DROID-SLAM [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>Mono/stereo/RGB-D (images; optional depth)</td>
</tr>
<tr>
<td>Dense RGB-D mapping</td>
<td>BundleFusion [<xref ref-type="bibr" rid="ref-26">26</xref>], NICE-SLAM [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>RGB-D, intrinsics</td>
</tr>
<tr>
<td>Learned matching front-end</td>
<td>SuperPoint [<xref ref-type="bibr" rid="ref-28">28</xref>], LoFTR [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>Images</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Evaluation Protocol and Reporting</title>
<p>Each sequence is evaluated over <italic>R</italic> independent runs; mean <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> std is reported when applicable. Failure rate is <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mtext>Fail</mml:mtext><mml:mrow><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x0023;</mml:mi><mml:mtext>failed runs</mml:mtext></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0023;</mml:mi><mml:mtext>total runs</mml:mtext></mml:mrow></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, where a run is counted as failed if optimization becomes numerically invalid (e.g., NaN/Inf states or normal equations) or if tracking/trajectory estimation breaks down irrecoverably before a complete trajectory can be aligned to ground truth. Absolute trajectory error (ATE) uses a single global least-squares alignment with the standard Umeyama procedure; stereo, RGB-D, and visual-inertial settings use <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mrow><mml:mi mathvariant="normal">S</mml:mi><mml:mi mathvariant="normal">E</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> alignment, whereas monocular settings use <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mrow><mml:mi mathvariant="normal">S</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>3</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> when scale is not observable. Key settings are listed in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Evaluation settings used throughout (unless otherwise stated). &#x201C;GN&#x201D; denotes Gauss-Newton, &#x201C;Seqs&#x201D; denotes sequences, &#x201C;seq&#x201D; denotes sequence, &#x201C;iters&#x201D; denotes iterations, &#x201C;MH&#x201D; denotes Machine Hall, &#x201C;V&#x201D; denotes Vicon Room, &#x201C;Val&#x201D; denotes validation, and &#x201C;Sel.&#x201D; denotes selected.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Seqs</th>
<th>Runs/seq (<italic>R</italic>)</th>
<th>Resolution</th>
<th>Window (<italic>N</italic>)</th>
<th>GN Iters (<italic>K</italic>)</th>
</tr>
</thead>
<tbody>
<tr>
<td>KITTI</td>
<td>00&#x2013;10</td>
<td>5</td>
<td><inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mn>1241</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>376</mml:mn></mml:math></inline-formula></td>
<td>7</td>
<td>12</td>
</tr>
<tr>
<td>KITTI-360</td>
<td>00, 02-06, 09</td>
<td>5</td>
<td><inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mn>1408</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>376</mml:mn></mml:math></inline-formula></td>
<td>7</td>
<td>12</td>
</tr>
<tr>
<td>EuRoC</td>
<td>MH &#x002B; V (All)</td>
<td>10</td>
<td><inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mn>752</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>480</mml:mn></mml:math></inline-formula></td>
<td>10</td>
<td>10</td>
</tr>
<tr>
<td>TUM</td>
<td>Fr1/2/3 (Sel.)</td>
<td>5</td>
<td><inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mn>640</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>480</mml:mn></mml:math></inline-formula></td>
<td>10</td>
<td>10</td>
</tr>
<tr>
<td>7-Scenes</td>
<td>All 7</td>
<td>5</td>
<td><inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mn>640</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>480</mml:mn></mml:math></inline-formula></td>
<td>12</td>
<td>8</td>
</tr>
<tr>
<td>ScanNet</td>
<td>Val (10 seqs)</td>
<td>3</td>
<td><inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mn>640</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>480</mml:mn></mml:math></inline-formula></td>
<td>12</td>
<td>8</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Reported runtime/frames per second (FPS) includes feature extraction, factor construction, windowed optimization, and the amortized loop-closure cost under the configured proposal/verification schedule. Per-stage timing for retrieval and verification is additionally logged to support reproducibility.</p>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Metrics</title>
<p>Absolute trajectory error (ATE)/relative pose error (RPE) [<xref ref-type="bibr" rid="ref-20">20</xref>], loop-closure and relocalization precision/recall, mapping quality where depth exists, uncertainty metrics, and efficiency (FPS/memory) are reported. For uncertainty, let <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula> be the pose error and <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mi mathvariant="normal">&#x03A3;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> the posterior covariance of the pose block at the solved state, approximated by the corresponding block of the inverse GN/LM normal matrix (local Laplace approximation, not an exact posterior). Assuming Gaussian errors, the average negative log-likelihood (NLL) is</p>
<p><disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:mtext>NLL</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>T</mml:mi></mml:mfrac><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mi>t</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi mathvariant="normal">&#x03A3;</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03A3;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mo>,</mml:mo></mml:math></disp-formula>with constant <italic>C</italic> shared across methods. Expected calibration error (ECE) follows standard binning of nominal vs. empirical coverage; reliability curves and score distributions are reported in <xref ref-type="sec" rid="s7_5">Section 7.5</xref>. False loops per kilometer use traveled distance along the reference trajectory.</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Results and Analysis</title>
<sec id="s7_1">
<label>7.1</label>
<title>Outdoor Odometry (KITTI/KITTI-360)</title>
<p>Outdoor trajectory estimation is evaluated under <xref ref-type="sec" rid="s6_2">Section 6.2</xref>. Results are summarized in <xref ref-type="table" rid="table-4">Table 4</xref>, and representative trajectories are visualized in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. On KITTI-360, <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula> shows reduced drift trends under broader appearance variation, aligning with conservative verified loop insertion and moderated factor influence via online covariance rescaling.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Outdoor trajectory accuracy on KITTI Odometry (averaged over sequences 00&#x2013;10) and KITTI-360. Monocular uses Sim(3) alignment; stereo uses SE(3). Absolute trajectory error (ATE) root-mean-square error (RMSE) [m]; relative pose error (RPE): translation [%]/rotation [&#x00B0;/100 m]; runtime is reported in frames per second (FPS).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Modality</th>
<th>KITTI ATE<inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>KITTI RPE-t<inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>KITTI RPE-r<inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>KITTI-360 ATE<inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Runtime (FPS)<inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>ORB-SLAM3 [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>Stereo</td>
<td>0.82 <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.15</td>
<td>0.95</td>
<td>0.28</td>
<td>3.42 <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.1</td>
<td>32.5</td>
</tr>
<tr>
<td>DSO [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>Mono</td>
<td>2.51 <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.85</td>
<td>2.15</td>
<td>0.54</td>
<td>12.8 <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 4.2</td>
<td>28.0</td>
</tr>
<tr>
<td>DROID-SLAM [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>Stereo</td>
<td>0.65 <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.12</td>
<td>0.78</td>
<td>0.21</td>
<td>2.15 <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.6</td>
<td>14.5</td>
</tr>
<tr>
<td><inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td>Mono</td>
<td>1.05 <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.22</td>
<td>1.25</td>
<td>0.38</td>
<td>3.25 <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 1.1</td>
<td>18.2</td>
</tr>
<tr>
<td><inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td>Stereo</td>
<td>0.64 <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.09</td>
<td>0.75</td>
<td>0.19</td>
<td>2.12 <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mo>&#x00B1;</mml:mo></mml:math></inline-formula> 0.5</td>
<td>17.5</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Qualitative trajectories on (<bold>a</bold>) KITTI-360 Sequence 00, (<bold>b</bold>) KITTI Odometry Sequence 02, and (<bold>c</bold>) KITTI Odometry Sequence 05. Alignment follows <xref ref-type="sec" rid="s6_2">Section 6.2</xref>.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-5.tif"/>
</fig>
</sec>
<sec id="s7_2">
<label>7.2</label>
<title>Visual-Inertial SLAM (EuRoC MAV)</title>
<p>On EuRoC with stereo &#x002B; IMU, <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula> attains errors comparable to established VIO baselines (<xref ref-type="table" rid="table-5">Table 5</xref>) and remains stable across runs. Calibration primarily acts as a moderation mechanism when visual residuals become temporarily unreliable (e.g., motion blur), reducing their dominance relative to inertial constraints.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>EuRoC Micro Aerial Vehicle (MAV) visual-inertial results. Absolute trajectory error (ATE) root-mean-square error (RMSE) [m] is averaged over Machine Hall (MH 01&#x2013;05) and Vicon Room (V1&#x2013;V2). &#x201C;IMU&#x201D; denotes inertial measurement unit, &#x201C;Fail%&#x201D; denotes failure rate computed over total runs (<inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn>50</mml:mn></mml:math></inline-formula> for MH, <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn>20</mml:mn></mml:math></inline-formula> for V), and runtime is reported in frames per second (FPS).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Sensors</th>
<th>ATE (MH)<inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Fail% (MH)<inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>ATE (V)<inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Fail% (V)<inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Runtime (FPS)<inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>OKVIS [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>Stereo &#x002B; IMU</td>
<td>0.052</td>
<td>0.0%</td>
<td>0.085</td>
<td>5.0%</td>
<td>25.0</td>
</tr>
<tr>
<td>VINS-Fusion</td>
<td>Stereo &#x002B; IMU</td>
<td>0.044</td>
<td>0.0%</td>
<td>0.038</td>
<td>5.0%</td>
<td>22.0</td>
</tr>
<tr>
<td>ORB-SLAM3 [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>Stereo &#x002B; IMU</td>
<td>0.033</td>
<td>0.0%</td>
<td>0.041</td>
<td>0.0%</td>
<td>24.5</td>
</tr>
<tr>
<td><inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td>Stereo &#x002B; IMU</td>
<td>0.034</td>
<td>0.0%</td>
<td>0.037</td>
<td>0.0%</td>
<td>16.8</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s7_3">
<label>7.3</label>
<title>Indoor RGB-D Tracking and Auxiliary Mapping Evidence (TUM, ScanNet, 7-Scenes)</title>
<p>Indoor RGB-D scenes contain textureless regions and perceptual aliasing. Across TUM, ScanNet, and 7-Scenes (<xref ref-type="table" rid="table-6">Table 6</xref>), <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula> achieves competitive tracking accuracy and strong relocalization, indicating that foundation features and verified loop insertion help maintain consistency under ambiguity. Because the main contribution of this work is conformal factor calibration rather than dense reconstruction itself, the ScanNet mesh results below are treated as auxiliary evidence that better-calibrated factor weighting also preserves downstream geometric consistency, rather than as the primary proof of the method.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Indoor trajectory accuracy on TUM red-green-blue-depth (RGB-D), ScanNet, and 7-Scenes. Values are absolute trajectory error (ATE) root-mean-square error (RMSE) [cm], runtime is reported in frames per second (FPS), and &#x201C;Reloc.&#x201D; denotes relocalization.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Input</th>
<th>TUM ATE<inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>ScanNet ATE<inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>7-Scenes ATE<inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Reloc. Rate<inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>FPS<inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>ORB-SLAM3 [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>RGB-D</td>
<td>1.8</td>
<td>7.2</td>
<td>4.5</td>
<td>82.5%</td>
<td>30.0</td>
</tr>
<tr>
<td>BundleFusion [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>RGB-D</td>
<td>2.5</td>
<td>8.5</td>
<td>5.1</td>
<td>&#x2013;</td>
<td>10.5</td>
</tr>
<tr>
<td>NICE-SLAM [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>RGB-D</td>
<td>2.1</td>
<td>9.8</td>
<td>6.5</td>
<td>&#x2013;</td>
<td>0.8</td>
</tr>
<tr>
<td>DROID-SLAM [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>RGB-D</td>
<td>2.0</td>
<td>6.4</td>
<td>3.8</td>
<td>91.0%</td>
<td>12.5</td>
</tr>
<tr>
<td><inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td>RGB-D</td>
<td>1.9</td>
<td>6.3</td>
<td>3.75</td>
<td>92.5%</td>
<td>16.0</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As auxiliary evidence of graph consistency on RGB-D data, quantitative dense mapping results on ScanNet are reported in <xref ref-type="table" rid="table-7">Table 7</xref>, and qualitative examples are shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Dense mapping quality on ScanNet (ground-truth mesh). Accuracy (Acc.) [cm] is distance to reference mesh; Completeness (Comp.) [%] uses a 5 cm threshold; Chamfer is in cm; peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) are also reported.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Dataset</th>
<th>Acc. (cm)<inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Comp. (%)<inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>Chamfer<inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>PSNR<inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>SSIM<inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>BundleFusion [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>ScanNet</td>
<td>6.8</td>
<td>72.1</td>
<td>7.5</td>
<td>20.1</td>
<td>0.78</td>
</tr>
<tr>
<td><inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td>ScanNet</td>
<td>6.5</td>
<td>74.5</td>
<td>7.2</td>
<td>24.2</td>
<td>0.84</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Qualitative dense reconstruction on ScanNet. Top: reconstructed meshes. Bottom: distance-to-mesh error heatmaps (shared scale).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-6.tif"/>
</fig>
</sec>
<sec id="s7_4">
<label>7.4</label>
<title>Loop Closure and Relocalization</title>
<p><xref ref-type="table" rid="table-8">Table 8</xref> reports loop closure precision/recall and relocalization. Geometric verification filters most spurious retrieval candidates, while descriptor-based proposal improves recall under large viewpoint/appearance changes. To directly isolate the necessity of the second stage, an additional comparison between descriptor-only loop closure and descriptor retrieval followed by geometric verification is reported in <xref ref-type="table" rid="table-9">Table 9</xref>. The verification stage is intended as a conservative gate before loop-factor insertion, since even a small number of false loop constraints can bias subsequent graph optimization. Qualitative verified examples are shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Loop closure (LC) and relocalization metrics computed on loop factors inserted after geometric verification. &#x201C;Prec.&#x201D; denotes precision, &#x201C;Reloc@<inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mi>d</mml:mi></mml:math></inline-formula>&#x201D; denotes relocalization within distance threshold <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mi>d</mml:mi></mml:math></inline-formula>, and False LC/km denotes false loop closures per traveled kilometer.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Dataset</th>
<th>LC Prec.<inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>LC Recall<inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>Reloc&#x0040;5 cm<inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>Reloc&#x0040;10 cm<inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>False LC/km<inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>ORB-SLAM3 [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>KITTI-360</td>
<td>100%</td>
<td>45.2%</td>
<td>65.5%</td>
<td>78.2%</td>
<td>0.00</td>
</tr>
<tr>
<td>DROID-SLAM [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>KITTI-360</td>
<td>95.5%</td>
<td>72.1%</td>
<td>82.1%</td>
<td>88.5%</td>
<td>0.05</td>
</tr>
<tr>
<td><inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td>KITTI-360</td>
<td>99.5%</td>
<td>73.2%</td>
<td>83.5%</td>
<td>89.8%</td>
<td>0.04</td>
</tr>
<tr>
<td>ORB-SLAM3 [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>TUM Loop</td>
<td>100%</td>
<td>52.8%</td>
<td>58.4%</td>
<td>70.1%</td>
<td>0.00</td>
</tr>
<tr>
<td><inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula></td>
<td>TUM Loop</td>
<td>99.8%</td>
<td>54.5%</td>
<td>60.2%</td>
<td>72.5%</td>
<td>0.01</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Direct ablation of descriptor-only loop closure vs. descriptor retrieval &#x002B; geometric verification on KITTI-360. &#x201C;LC&#x201D; denotes loop closure, absolute trajectory error (ATE) is reported in meters, and runtime is reported in frames per second (FPS).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Setting</th>
<th>LC Prec.<inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>LC Recall<inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th>False LC/km<inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>ATE [m]<inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>FPS<inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>Descriptor-only</td>
<td>94.1%</td>
<td>75.6%</td>
<td>0.36</td>
<td>3.38</td>
<td>18.4</td>
</tr>
<tr>
<td>Descriptor &#x002B; geometric verification (<inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula>)</td>
<td>99.5%</td>
<td>73.2%</td>
<td>0.04</td>
<td>2.12</td>
<td>17.5</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Loop-closure examples passing geometric verification: query/retrieved frames, verified correspondences, and inserted loop constraint.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-7.tif"/>
</fig>
</sec>
<sec id="s7_5">
<label>7.5</label>
<title>Uncertainty, Robustness, and Cross-Sensor Shift</title>
<p>Uncertainty is evaluated with negative log-likelihood (NLL)/expected calibration error (ECE) and robustness with failure rate under zero-shot transfer, with emphasis on how conformal calibration behaves under cross-dataset and cross-sensor shift. The transfer from KITTI to KITTI-360 mainly changes appearance statistics, scene layout, and long-horizon driving context under the same stereo sensing regime, whereas the transfer from TartanAir to EuRoC additionally introduces synthetic-to-real, motion-regime, and stereo &#x002B; IMU VIO differences. In fixed-noise SLAM systems, such shifts can mis-scale visual, depth, and inertial factor families, causing some modalities to dominate the normal equations. The learned heads provide modality-conditioned initial covariances, and the conformal layer further corrects residual scale mismatch online per factor family. <xref ref-type="table" rid="table-10">Table 10</xref> and <xref ref-type="fig" rid="fig-8">Fig. 8</xref> show that conformal calibration reduces ECE and is accompanied by lower failure rates across both transfer settings, consistent with more reliable factor weighting under shift rather than merely better in-domain fitting.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Cross-dataset and cross-sensor uncertainty calibration and robustness under zero-shot transfer. Negative log-likelihood (NLL), expected calibration error (ECE), and absolute trajectory error (ATE) follow <xref ref-type="sec" rid="s6_3">Section 6.3</xref>; &#x201C;Fail%&#x201D; denotes failure rate and &#x201C;Med.&#x201D; denotes median.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Setting</th>
<th>Transfer</th>
<th>NLL<inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>ECE<inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Fail%<inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Med. ATE [m]<inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>95% ATE [m]<inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>Uncalibrated</td>
<td>KITTI <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> KITTI-360</td>
<td>2.45</td>
<td>0.32</td>
<td>2.9%</td>
<td>2.20</td>
<td>2.85</td>
</tr>
<tr>
<td>Conformal</td>
<td>KITTI <inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> KITTI-360</td>
<td>2.12</td>
<td>0.06</td>
<td>0.0%</td>
<td>2.12</td>
<td>2.65</td>
</tr>
<tr>
<td>Uncalibrated</td>
<td>TartanAir <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> EuRoC</td>
<td>1.80</td>
<td>0.41</td>
<td>5.5%</td>
<td>0.055</td>
<td>0.095</td>
</tr>
<tr>
<td>Conformal</td>
<td>TartanAir <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> EuRoC</td>
<td>1.40</td>
<td>0.10</td>
<td>1.8%</td>
<td>0.037</td>
<td>0.082</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Uncertainty calibration diagnostics. Top: reliability diagrams. Bottom: score distributions under transfer.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-8a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-8b.tif"/>
</fig>
</sec>
</sec>
<sec id="s8">
<label>8</label>
<title>Ablations</title>
<p>Key components are ablated under the same protocol, with particular emphasis on calibration under shift and on the loop-closure design. <xref ref-type="table" rid="table-11">Table 11</xref> shows that removing conformal calibration increases both ATE and failure rate (A1 vs. A0), replacing the foundation backbone degrades transfer robustness (A2), and geometric-only loop closure substantially increases drift (A3), indicating the value of descriptor-topological proposal. Together with the cross-shift results in <xref ref-type="table" rid="table-10">Table 10</xref>, these ablations support the claim that conformal reweighting is the primary mechanism improving robustness under transfer. <xref ref-type="table" rid="table-9">Table 9</xref> further isolates the contribution of the geometric verification stage beyond descriptor retrieval alone.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Module ablation on KITTI-360 (Validation). A0: full method. A2 replaces DINOv2 with ResNet50. &#x201C;LC&#x201D; denotes loop closure, &#x201C;ATE&#x201D; denotes absolute trajectory error, &#x201C;calib.&#x201D; denotes calibration, &#x201C;w/o&#x201D; denotes without, and &#x201C;descriptor-topo&#x201D; denotes descriptor-topological, and Fail% corresponds to integer failures (1/35 <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mo>&#x2248;</mml:mo></mml:math></inline-formula> 2.9%).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Variant</th>
<th>Backbone</th>
<th>Conformal calib.</th>
<th>LC factor</th>
<th>Toggles</th>
<th>ATE [m]<inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th>Fail%<inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>A0 (full)</td>
<td>DINOv2</td>
<td>&#x2713;</td>
<td>Descriptor-topo</td>
<td>Full</td>
<td>2.12</td>
<td>0.0%</td>
</tr>
<tr>
<td>A1</td>
<td>DINOv2</td>
<td>&#x2013;</td>
<td>Descriptor-topo</td>
<td>Full</td>
<td>2.20</td>
<td>2.9%</td>
</tr>
<tr>
<td>A2</td>
<td>ResNet50</td>
<td>&#x2713;</td>
<td>Descriptor-topo</td>
<td>Full</td>
<td>3.12</td>
<td>5.7%</td>
</tr>
<tr>
<td>A3</td>
<td>DINOv2</td>
<td>&#x2713;</td>
<td>Geometric-only</td>
<td>Full</td>
<td>4.55</td>
<td>2.9%</td>
</tr>
<tr>
<td>A4</td>
<td>DINOv2</td>
<td>&#x2713;</td>
<td>Descriptor-topo</td>
<td>w/o depth prior</td>
<td>2.95</td>
<td>2.9%</td>
</tr>
<tr>
<td>A5</td>
<td>DINOv2</td>
<td>&#x2713;</td>
<td>Descriptor-topo</td>
<td>w/o motion prior</td>
<td>2.25</td>
<td>2.9%</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s9">
<label>9</label>
<title>Discussion and Limitations</title>
<p>Solver stability in SLAM is tightly coupled to factor weighting. In this work, &#x201C;stability&#x201D; is used in an operational sense: fewer failed runs, fewer catastrophic drifts/divergences, and less sensitivity to misweighted factor families during sequential GN/LM updates under transfer. Under domain shift, miscalibrated uncertainties can overweight unreliable constraints and degrade conditioning, producing drift or divergence. The conformal rescaling rule in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref> provides a simple online mechanism to adjust covariance magnitudes using observed residual statistics, without retraining, and the empirical evidence in <xref ref-type="table" rid="table-10">Tables 10</xref> and <xref ref-type="table" rid="table-11">11</xref> should be interpreted in this operational sense rather than as a stand-alone spectral conditioning proof. A representative divergence-vs.-recovery comparison is shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Failure case analysis on KITTI<inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>KITTI-360. Row 1: divergence without calibration. Row 2: recovery with conformal covariance rescaling.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79663-fig-9.tif"/>
</fig>
<p>Limitations include: (i) foundation backbones increase compute relative to compact convolutional neural network (CNN) front-ends, motivating distillation or reduced token resolution for real-time deployment; (ii) calibration relies on bounded windows and approximate stationarity, so abrupt regime changes can challenge residual statistics despite warm-up and caps; (iii) the conformal-style update is empirical and does not provide strict exchangeability-based coverage guarantees in sequential SLAM; and (iv) loop closure remains subject to the retrieval/verification trade-off under severe viewpoint changes and repeated structures. In highly dynamic scenes with large non-rigid occluders, retrieval and correspondences may degrade; integrating motion- or instance-aware masking with dynamic-aware matching is a natural extension that preserves the current solver formulation. The multi-stage loop-closure pipeline also introduces controllable overhead; sustained real-time operation depends on scheduling choices such as keyframe triggering, candidate caps, and (when available) asynchronous verification.</p>
</sec>
<sec id="s10">
<label>10</label>
<title>Conclusion</title>
<p>This article introduced <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">F</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mtext>-</mml:mtext><mml:mi>S</mml:mi><mml:mi>L</mml:mi><mml:mi>A</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:math></inline-formula>, a sensor-agnostic SLAM framework combining frozen foundation representations, probabilistic factor heads, and a classical factor-graph back-end equipped with online conformal calibration. By rescaling factor covariances using residual quantiles per factor type, the method specifically targets systematic misweighting under cross-dataset and cross-sensor shift and yields more reliable optimization behavior under transfer, reflected in lower failure rates and reduced sensitivity to misweighted factors. Descriptor-topological loop closure was also described based on foundation descriptors with geometric verification for conservative loop insertion, and dataset sensor fields and evaluation protocols were documented for reproducible cross-modal comparison.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The author received no specific funding for this study.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The datasets analyzed in this study are publicly available:</p>
<p>&#x2022; &#x2002;&#x2002;KITTI Odometry: <ext-link ext-link-type="uri" xlink:href="https://www.cvlibs.net/datasets/kitti/eval_odometry.php">https://www.cvlibs.net/datasets/kitti/eval_odometry.php</ext-link></p>
<p>&#x2022; &#x2002;&#x2002;KITTI-360: <ext-link ext-link-type="uri" xlink:href="https://www.cvlibs.net/datasets/kitti-360/">https://www.cvlibs.net/datasets/kitti-360/</ext-link></p>
<p>&#x2022; &#x2002;&#x2002;EuRoC MAV: <ext-link ext-link-type="uri" xlink:href="https://ethz-asl.github.io/datasets/">https://ethz-asl.github.io/datasets/</ext-link></p>
<p>&#x2022; &#x2002;&#x2002;TUM RGB-D: <ext-link ext-link-type="uri" xlink:href="https://cvg.cit.tum.de/data/datasets/rgbd-dataset">https://cvg.cit.tum.de/data/datasets/rgbd-dataset</ext-link></p>
<p>&#x2022; &#x2002;&#x2002;ScanNet: <ext-link ext-link-type="uri" xlink:href="https://github.com/ScanNet/ScanNet">https://github.com/ScanNet/ScanNet</ext-link></p>
<p>&#x2022; &#x2002;&#x2002;7-Scenes: <ext-link ext-link-type="uri" xlink:href="https://www.microsoft.com/en-us/research/project/rgb-d-dataset-7-scenes/">https://www.microsoft.com/en-us/research/project/rgb-d-dataset-7-scenes/</ext-link></p>
<p>&#x2022; &#x2002;&#x2002;TartanAir: <ext-link ext-link-type="uri" xlink:href="https://theairlab.org/tartanair-dataset/">https://theairlab.org/tartanair-dataset/</ext-link></p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The author declares no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yeung</surname> <given-names>SK</given-names></string-name></person-group>. <article-title>Photo-SLAM: real-time simultaneous localization and photorealistic mapping for monocular stereo and RGB-D cameras</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2024</year>. p. <fpage>21584</fpage>&#x2013;<lpage>93</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Han</surname> <given-names>K</given-names></string-name></person-group>. <article-title>IBD-SLAM: learning image-based depth fusion for generalizable SLAM</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2024</year>. p. <fpage>10563</fpage>&#x2013;<lpage>73</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Oquab</surname> <given-names>M</given-names></string-name>, <string-name><surname>Darcet</surname> <given-names>T</given-names></string-name>, <string-name><surname>Moutakanni</surname> <given-names>T</given-names></string-name>, <string-name><surname>Vo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Szafraniec</surname> <given-names>M</given-names></string-name>, <string-name><surname>Khalidov</surname> <given-names>V</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DINOv2: learning robust visual features without supervision</article-title>. <comment>arXiv:2304.07193. 2023</comment>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abdelkarim</surname> <given-names>A</given-names></string-name>, <string-name><surname>Voos</surname> <given-names>H</given-names></string-name>, <string-name><surname>G&#x00F6;rges</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Factor graphs in optimization-based robotic control&#x2014;a tutorial and review</article-title>. <source>IEEE Access</source>. <year>2025</year>;<volume>13</volume>(<issue>23</issue>):<fpage>28315</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2025.3534993</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gibbs</surname> <given-names>I</given-names></string-name>, <string-name><surname>Cand&#x00E8;s</surname> <given-names>EJ</given-names></string-name></person-group>. <article-title>Conformal inference for online prediction with arbitrary distribution shifts</article-title>. <source>J Mach Learn Res</source>. <year>2024</year>;<volume>25</volume>(<issue>162</issue>):<fpage>1</fpage>&#x2013;<lpage>36</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Campos</surname> <given-names>C</given-names></string-name>, <string-name><surname>Elvira</surname> <given-names>R</given-names></string-name>, <string-name><surname>Rodr&#x00ED;guez</surname> <given-names>JJG</given-names></string-name>, <string-name><surname>Montiel</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Tard&#x00F3;s</surname> <given-names>JD</given-names></string-name></person-group>. <article-title>ORB-SLAM3: an accurate open-source library for visual, visual-inertial, and multimap SLAM</article-title>. <source>IEEE Trans Robot</source>. <year>2021</year>;<volume>37</volume>(<issue>6</issue>):<fpage>1874</fpage>&#x2013;<lpage>90</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lindenberger</surname> <given-names>P</given-names></string-name>, <string-name><surname>Sarlin</surname> <given-names>PE</given-names></string-name>, <string-name><surname>Pollefeys</surname> <given-names>M</given-names></string-name></person-group>. <article-title>LightGlue: local feature matching at light speed</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2023</year>. p. <fpage>17627</fpage>&#x2013;<lpage>38</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Edstedt</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Q</given-names></string-name>, <string-name><surname>B&#x00F6;kman</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wadenb&#x00E4;ck</surname> <given-names>M</given-names></string-name>, <string-name><surname>Felsberg</surname> <given-names>M</given-names></string-name></person-group>. <article-title>RoMa: robust dense feature matching</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2024</year>. p. <fpage>19790</fpage>&#x2013;<lpage>800</lpage>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Matsuki</surname> <given-names>H</given-names></string-name>, <string-name><surname>Murai</surname> <given-names>R</given-names></string-name>, <string-name><surname>Kelly</surname> <given-names>PHJ</given-names></string-name>, <string-name><surname>Davison</surname> <given-names>AJ</given-names></string-name></person-group>. <article-title>Gaussian splatting SLAM</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2024</year>. p. <fpage>18039</fpage>&#x2013;<lpage>48</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Keetha</surname> <given-names>N</given-names></string-name>, <string-name><surname>Karhade</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jatavallabhula</surname> <given-names>KM</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Scherer</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ramanan</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SplaTAM: splat track &#x0026; map 3D Gaussians for dense RGB-D SLAM</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2024</year>. p. <fpage>21357</fpage>&#x2013;<lpage>66</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Clart&#x00E9;</surname> <given-names>L</given-names></string-name>, <string-name><surname>Loureiro</surname> <given-names>B</given-names></string-name>, <string-name><surname>Krzakala</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zdeborov&#x00E1;</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Expectation consistency for calibration of neural networks</article-title>. In: <conf-name>Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence. Vol. 216. London, UK: PMLR</conf-name>; <year>2023</year>. p. <fpage>443</fpage>&#x2013;<lpage>53</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Oliveira</surname> <given-names>RI</given-names></string-name>, <string-name><surname>Orenstein</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ramos</surname> <given-names>T</given-names></string-name>, <string-name><surname>Romano</surname> <given-names>JV</given-names></string-name></person-group>. <article-title>Split conformal prediction and non-exchangeable data</article-title>. <source>J Mach Learn Res</source>. <year>2024</year>;<volume>25</volume>(<issue>225</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xue</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Resilient routing: risk-aware dynamic routing in smart logistics via spatiotemporal graph learning</article-title>. <comment>arXiv:2601.13632. 2026</comment>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qin</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>S</given-names></string-name></person-group>. <article-title>VINS-Mono: a robust and versatile monocular visual-inertial state estimator</article-title>. <source>IEEE Trans Robot</source>. <year>2018</year>;<volume>34</volume>(<issue>4</issue>):<fpage>1004</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>X</given-names></string-name></person-group>. <article-title>LoFTR: detector-free local feature matching with transformers</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2021</year>. p. <fpage>8922</fpage>&#x2013;<lpage>31</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Geiger</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lenz</surname> <given-names>P</given-names></string-name>, <string-name><surname>Urtasun</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Are we ready for autonomous driving? The KITTI vision benchmark suite</article-title>. In: <conf-name>2012 IEEE Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-name>Piscataway, NJ, USA: IEEE</publisher-name>; <year>2012</year>. p. <fpage>3354</fpage>&#x2013;<lpage>61</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Geiger</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lenz</surname> <given-names>P</given-names></string-name>, <string-name><surname>Stiller</surname> <given-names>C</given-names></string-name>, <string-name><surname>Urtasun</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Vision meets robotics: the KITTI dataset</article-title>. <source>Int J Robot Res</source>. <year>2013</year>;<volume>32</volume>(<issue>11</issue>):<fpage>1231</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Geiger</surname> <given-names>A</given-names></string-name></person-group>. <article-title>KITTI-360: a novel dataset and benchmarks for urban scene understanding in 2D and 3D</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2022</year>;<volume>45</volume>(<issue>3</issue>):<fpage>3292</fpage>&#x2013;<lpage>310</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Burri</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nikolic</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gohl</surname> <given-names>P</given-names></string-name>, <string-name><surname>Schneider</surname> <given-names>T</given-names></string-name>, <string-name><surname>Rehder</surname> <given-names>J</given-names></string-name>, <string-name><surname>Omari</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The EuRoC micro aerial vehicle datasets</article-title>. <source>Int J Robot Res</source>. <year>2016</year>;<volume>35</volume>(<issue>10</issue>):<fpage>1157</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1177/0278364915620033</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sturm</surname> <given-names>J</given-names></string-name>, <string-name><surname>Engelhard</surname> <given-names>N</given-names></string-name>, <string-name><surname>Endres</surname> <given-names>F</given-names></string-name>, <string-name><surname>Burgard</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cremers</surname> <given-names>D</given-names></string-name></person-group>. <article-title>A benchmark for the evaluation of RGB-D SLAM systems</article-title>. In: <conf-name>2012 IEEE/RSJ International Conference on Intelligent Robots and Systems</conf-name>. <publisher-name>Piscataway, NJ, USA: 	IEEE</publisher-name>; <year>2012</year>. p. <fpage>573</fpage>&#x2013;<lpage>80</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dai</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>AX</given-names></string-name>, <string-name><surname>Savva</surname> <given-names>M</given-names></string-name>, <string-name><surname>Halber</surname> <given-names>M</given-names></string-name>, <string-name><surname>Funkhouser</surname> <given-names>T</given-names></string-name>, <string-name><surname>Nie&#x00DF;ner</surname> <given-names>M</given-names></string-name></person-group>. <article-title>ScanNet: richly-annotated 3D reconstructions of indoor scenes</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2017</year>. p. <fpage>5828</fpage>&#x2013;<lpage>39</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Engel</surname> <given-names>J</given-names></string-name>, <string-name><surname>Koltun</surname> <given-names>V</given-names></string-name>, <string-name><surname>Cremers</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Direct sparse odometry</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2017</year>;<volume>40</volume>(<issue>3</issue>):<fpage>611</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2017.2658577</pub-id>; <pub-id pub-id-type="pmid">28422651</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Forster</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gassner</surname> <given-names>M</given-names></string-name>, <string-name><surname>Werlberger</surname> <given-names>M</given-names></string-name>, <string-name><surname>Scaramuzza</surname> <given-names>D</given-names></string-name></person-group>. <article-title>SVO: semidirect visual odometry for monocular and multicamera systems</article-title>. <source>IEEE Trans Robot</source>. <year>2016</year>;<volume>33</volume>(<issue>2</issue>):<fpage>249</fpage>&#x2013;<lpage>65</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Leutenegger</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lynen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bosse</surname> <given-names>M</given-names></string-name>, <string-name><surname>Siegwart</surname> <given-names>R</given-names></string-name>, <string-name><surname>Furgale</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Keyframe-based visual-inertial odometry using nonlinear optimization</article-title>. <source>Int J Robot Res</source>. <year>2015</year>;<volume>34</volume>(<issue>3</issue>):<fpage>314</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1177/0278364914554813</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Teed</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>J</given-names></string-name></person-group>. <article-title>DROID-SLAM: deep visual slam for monocular, stereo, and RGB-D cameras</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2021</year>;<volume>34</volume>:<fpage>16558</fpage>&#x2013;<lpage>69</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dai</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nie&#x00DF;ner</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zollh&#x00F6;fer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Izadi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Theobalt</surname> <given-names>C</given-names></string-name></person-group>. <article-title>BundleFusion: real-time globally consistent 3D reconstruction using on-the-fly surface reintegration</article-title>. <source>ACM Trans Graph</source>. <year>2017</year>;<volume>36</volume>(<issue>4</issue>):<fpage>1</fpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Larsson</surname> <given-names>V</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Bao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Nice-SLAM: neural implicit scalable encoding for SLAM</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2022</year>. p. <fpage>12786</fpage>&#x2013;<lpage>96</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>DeTone</surname> <given-names>D</given-names></string-name>, <string-name><surname>Malisiewicz</surname> <given-names>T</given-names></string-name>, <string-name><surname>Rabinovich</surname> <given-names>A</given-names></string-name></person-group>. <article-title>SuperPoint: self-supervised interest point detection and description</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops</conf-name>. <publisher-loc>Piscataway, NJ, USA: IEEE</publisher-loc>; <year>2018</year>. p. <fpage>224</fpage>&#x2013;<lpage>36</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>
