Using PCA and K-Means Clustering for Ink Discrimination in Hyperspectral Document Imaging

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract This paper presents an analysis of a hyperspectral image (HSI) of a document with the goal of detecting different inks. The approach includes an analysis of the spectral bands, visual and spectral analysis of the text lines, dimensionality reduction using Principal Component Analysis (PCA), and unsupervised clustering using K-Means to estimate ink categories. The methodology is validated through a comparison of clustering before and after PCA. Visual outputs and spectral plots are provided to display observed trends and ink separation results.
Full text 30,696 characters · extracted from preprint-html · click to expand
Using PCA and K-Means Clustering for Ink Discrimination in Hyperspectral Document Imaging | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Short Report Using PCA and K-Means Clustering for Ink Discrimination in Hyperspectral Document Imaging Maryam Batool This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7129735/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This paper presents an analysis of a hyperspectral image (HSI) of a document with the goal of detecting different inks. The approach includes an analysis of the spectral bands, visual and spectral analysis of the text lines, dimensionality reduction using Principal Component Analysis (PCA), and unsupervised clustering using K-Means to estimate ink categories. The methodology is validated through a comparison of clustering before and after PCA. Visual outputs and spectral plots are provided to display observed trends and ink separation results. Theoretical Computer Science Hyperspectral Images Principal Component Analysis K-Means Figures Figure 1 Figure 2 Figure 3 Figure 4 I. Introduction Hyperspectral imaging provides high-resolution spectral data for each pixel, making it a powerful non-destructive method for document analysis[1]. Detecting different inks is valuable in forgery detection and heritage document authentication. Previous works have explored both statistical and machine learning approaches for ink detection[2][3]. Dimensionality reduction using PCA is widely adopted for visual analysis and reducing computational cost, while clustering methods such as K-Means are effective for unsupervised pattern recognition. II. LITERATURE REVIEW Hyperspectral Imaging (HSI) has emerged as a transformative technology across diverse domains—ranging from agriculture and medicine to defense and forensic sciences—due to its ability to capture rich spectral and spatial information beyond the human visual range [1]. In forensic document examination, where accuracy and document preservation are paramount, HSI enables non-destructive analysis of inks that conventional techniques often fail to differentiate. A pivotal contribution by Z. Khan et al. [2] demonstrated that HSI could successfully identify ink mismatches in a non-invasive manner, preserving the physical integrity of questioned documents. This foundational work was further extended by Abbas et al. [3], who proposed an automated pipeline leveraging HSI's spectral sensitivity to detect ink discrepancies—bringing forensic analysis closer to real-time, scalable solutions. Earlier innovations, such as those by Ostrum [4], and later expanded upon by Khan et al. [5], laid the groundwork for hyperspectral-based ink analysis. Ostrum highlighted the utility of tunable filters in detecting ink obliterations, while Khan and colleagues explored near-infrared hyperspectral imaging with Principal Component Analysis (PCA) to effectively identify document forgeries. Subsequent approaches adopted advanced dimensionality reduction techniques. Devassy and George [6] utilized t-distributed stochastic neighbor embedding (t-SNE) to visually cluster spectrally similar inks, aiding manual review. Addressing class imbalance—a common issue when certain inks are underrepresented—Nasir et al. [7] proposed hyperspectral unmixing methods to improve mismatch detection in such challenging conditions. The incorporation of deep learning has significantly advanced HSI capabilities. M. J. Khan et al. [8] pioneered the use of Convolutional Neural Networks (CNNs) for ink classification, achieving up to 98% accuracy in blue ink detection. Their subsequent research refined deep learning architectures for more complex scenarios, including black ink differentiation [9], establishing state-of-the-art benchmarks. To enhance efficiency without compromising accuracy, Uzair et al. [10] introduced joint sparse band selection techniques. These approaches reduced computational burden by selecting only the most discriminative spectral bands for ink detection, ensuring faster yet reliable processing. Complementary contributions such as those by Khan et al. [5] proposed fuzzy clustering and deep models for forgery detection, while a hybrid spatio-spectral CNN model was also introduced for document authentication [6]. Newer frameworks further enhance accuracy by combining classification and segmentation strategies. For instance, Saleem et al. [11] introduced a patch-based hybrid deep learning model that improves ink-region localization using spectral-spatial cues. Additionally, the move toward real-time applications has begun. Raza et al. [12] developed a portable HSI system integrated with lightweight models for in-field document verification, addressing deployment challenges. The field has also benefited from the creation of benchmark datasets, such as the iVision HHID by Ammad ul Islam et al. [8], which supports standardized evaluation of hyperspectral document analysis techniques. Further, broader reviews by Qureshi et al. [9] and Khan et al. [1] summarize current challenges, opportunities, and future directions in HSI applications. Together, these twelve studies chronicle the progression of hyperspectral imaging in document forensics—from early spectral sensing strategies to AI-powered, real-time analytical frameworks. The trajectory reflects a significant evolution toward scalable, automated, and non-invasive forensic tools rooted in spectral precision and machine intelligence. III. Methodology A. Band Analysis and Visualization (Part a and b) : We loaded an ENVI-format HSI cube using the Python Spectral library. The cube contains 224 spectral bands ranging approximately from 400 nm to 1000 nm. To analyze band content, we visualized the 1st, 30th, 60th, and last band as grayscale images. Visually, earlier bands are darker and highlight ink absorption. Mid-spectrum bands show greater text-background contrast, while higher bands fade, consistent with ink's low infrared reflectance as can be seen in Fig. 1. B. Spectral Response Analysis (Part c) : Binary masks were manually created for two lines of text. For each mask, average spectral reflectance was plotted across all bands. Line 1 and line 2 exhibit distinct spectral profiles, especially in the visible range (400-700nm), indicating different inks as shown in Fig. 2. C. Ink Detection via Clustering (Part d) : We applied K-Means clustering with k = 3 to the spectral pixels from the text masks. This unsupervised method grouped pixels based on spectral similarity. Strengths: Simple, interpretable, fast. Weaknesses: Assumes spherical clusters, sensitive to initialization, may underperform with overlapping spectra. D. Color-Labeling of Inks (Part e) : Cluster labels were mapped to colors to visualize ink regions. This shows clear separation between text lines with different inks as shown in Fig. 3. E. Dimensionality Reduction using PCA and Reapplying Clustering on PCA Output (Part f and g) : The HSI cube was reshaped and normalized, then PCA was applied to reduce 224 bands to 3 components. Each component was visualized and also used for clustering, the results for which can be observed in Fig. 4. K-Means was then reapplied on the PCA-transformed data. The results were comparable to the full-band method. Comparison of Clustering Methods : - K-Means (Full Bands): Input Shape N x 224 — Estimated Inks: 3 - K-Means (PCA): Input Shape N x 3 — Estimated Inks: IV. CONCLUSION The approach effectively identifies ink clusters using spectral features. PCA reduces computational load without sacrificing accuracy. K-Means offers a strong baseline method. Using all these methods we can analyze the image and identify the different inks that were used within the document in a non-destructive way. V. FUTURE WORK Future improvements include: - Applying spectral angle mapper or GMM for better clustering. - Integrating deep autoencoders for non-linear reduction. - Testing real-world scanned documents with noise. References Muhammad Jaleed Khan, Hamid Saeed Khan, Adeel Yousaf, Khurram Khurshid, Asad Abbas, “Modern Trends in Hyperspectral Image Analysis: A Review,” IEEE Access , 6(1):14118-14119, 2018. Z. Khan, F. Shafait, A. Mian, “Hyperspectral Imaging for Ink Mismatch Detection,” 12th Int’l Conf. on Document Analysis and Recognition (ICDAR) , 2013, pp. 877–881. A. Abbas, K. Khurshid, F. Shafait, “Towards Automated Ink Mismatch Detection in Hyperspectral Document Images,” 14th ICDAR , 2017, pp. 1229–1236. Ostrum, C. J., “Use of Tunable Filters in Ink Analysis,” Forensic Science International , Vol. 122, 2001, pp. 49–53. M. J. Khan, A. Yousaf, A. Abbas, K. Khurshid, “Deep Learning for Automated Forgery Detection in Hyperspectral Document Images,” Journal of Electronic Imaging , 27(5):053001, 2018. M. J. Khan, K. Khurshid, F. Shafait, “A Spatio-Spectral Hybrid Convolutional Architecture for Hyperspectral Document Authentication,” ICDAR , 2019, pp. 1097–1102. Nasir, A., Khan, M. J., Shafait, F., Khurshid, K., “Hyperspectral Unmixing for Ink Discrimination,” Digital Image Computing Techniques and Applications (DICTA) , 2017. Ammad ul Islam, M. J. Khan, M. Asad, H. A. Khan, K. Khurshid, “iVision HHID: Handwritten Hyperspectral Images Dataset for Benchmarking Hyperspectral Imaging-based Document Forensic Analysis,” Data in Brief , 2022. Rizwan Qureshi, Muhammad Uzair, Khurram Khurshid, Hong Yan, “Hyperspectral Document Image Processing: Applications, Challenges and Future Prospects,” Pattern Recognition , 90(1):12–22, 2019. M. Uzair, K. Khurshid, M. J. Khan, “Joint Sparse Band Selection for Hyperspectral Ink Analysis,” IEEE Transactions on Image Processing , vol. 29, pp. 4761–4773, 2020. Saleem, H., Yousaf, A., Khan, H. A., “Spectral-Spatial Ink Segmentation Using Hybrid Patch-Based Deep Learning for HSI Documents,” Forensic Science International: Reports , 2021. Raza, S., Khan, M. J., Khurshid, K., “Design and Deployment of a Portable Real-Time Hyperspectral Imaging System for Document Forensics,” Sensors , 2023, 23(2):455. Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7129735","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Short Report","associatedPublications":[],"authors":[{"id":485708405,"identity":"534c0f39-3109-4ea2-a7cd-1eefb1908ed4","order_by":0,"name":"Maryam Batool","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA70lEQVRIiWNgGAWjYDACCQY2GAMIKoCYmbmBFC1nQFoYSdHC2AZiEdCiO7uB7cHPHbXy8rMbWDf8nFcbzd8O1PKjYhtOLWZ3DrAb9p45brjhzgG2m73bjufOOMzYwNhz5jZuLTcS2CR4244xbpBIYLvBu+1YbgNQCzNjG34tkn/bjtnPn5HAdvPvnGO584nRIs3bVpPYAGTc5m2oyd1AUMudg+3Gsm0HkjfcSGy7LXPsQO5GoJaDeP1yu/nYw7dtdbbzZyQfu/mmpi533vnDBx/8qMCtBRoLh5EZDAwH8KiHgToMxigYBaNgFIwCOAAAZmVjFxPAKGkAAAAASUVORK5CYII=","orcid":"","institution":"","correspondingAuthor":true,"prefix":"","firstName":"Maryam","middleName":"","lastName":"Batool","suffix":""}],"badges":[],"createdAt":"2025-07-15 10:56:10","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-7129735/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7129735/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":87048156,"identity":"90790c10-6b33-4198-a26d-993e2dac3d6a","added_by":"auto","created_at":"2025-07-18 14:44:45","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":270022,"visible":true,"origin":"","legend":"\u003cp\u003e1st, 30th, 60th and the last band of the hyperspectral image as separate grayscale images.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7129735/v1/8c17c96d41c70a57c2fb9b80.png"},{"id":87048998,"identity":"06aae6e1-2156-4960-b1fa-9dc00ba4c735","added_by":"auto","created_at":"2025-07-18 14:52:45","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":81568,"visible":true,"origin":"","legend":"\u003cp\u003eSpectral Responses of Text lines\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7129735/v1/dfe33155c21c2141194b7d2f.png"},{"id":87048162,"identity":"26e429c2-21ab-4518-a550-24459736539c","added_by":"auto","created_at":"2025-07-18 14:44:45","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":165818,"visible":true,"origin":"","legend":"\u003cp\u003eClassifying text written with different inks in the document.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7129735/v1/3ecd737170dc568ee4302e5d.png"},{"id":87048160,"identity":"44c805a0-cb1f-4e6f-879a-7bc79780b434","added_by":"auto","created_at":"2025-07-18 14:44:45","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":643014,"visible":true,"origin":"","legend":"\u003cp\u003eDimensionality reduction of the original HSI cube using PCA.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7129735/v1/259f072f044e7fb0268224e1.png"},{"id":87053133,"identity":"bd3b4237-a959-4b2b-9db0-3bb1a1b1fb31","added_by":"auto","created_at":"2025-07-18 15:09:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1689863,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7129735/v1/cfac7a7e-ba37-42a9-97dd-b4ae0c248522.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eUsing PCA and K-Means Clustering for Ink Discrimination in Hyperspectral Document Imaging\u003c/p\u003e","fulltext":[{"header":"I. Introduction","content":"\u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eHyperspectral imaging provides high-resolution spectral data for each pixel, making it a powerful non-destructive method for document analysis[1]. Detecting different inks is valuable in forgery detection and heritage document authentication. Previous works have explored both statistical and machine learning approaches for ink detection[2][3]. Dimensionality reduction using PCA is widely adopted for visual analysis and reducing computational cost, while clustering methods such as K-Means are effective for unsupervised pattern recognition.\u003c/span\u003e\u003c/p\u003e"},{"header":"II.\tLITERATURE REVIEW","content":"\u003cp\u003eHyperspectral Imaging (HSI) has emerged as a transformative technology across diverse domains—ranging from agriculture and medicine to defense and forensic sciences—due to its ability to capture rich spectral and spatial information beyond the human visual range [1]. In forensic document examination, where accuracy and document preservation are paramount, HSI enables non-destructive analysis of inks that conventional techniques often fail to differentiate.\u003c/p\u003e\u003cp\u003eA pivotal contribution by Z. Khan et al. [2] demonstrated that HSI could successfully identify ink mismatches in a non-invasive manner, preserving the physical integrity of questioned documents. This foundational work was further extended by Abbas et al. [3], who proposed an automated pipeline leveraging HSI's spectral sensitivity to detect ink discrepancies—bringing forensic analysis closer to real-time, scalable solutions.\u003c/p\u003e\u003cp\u003eEarlier innovations, such as those by Ostrum [4], and later expanded upon by Khan et al. [5], laid the groundwork for hyperspectral-based ink analysis. Ostrum highlighted the utility of tunable filters in detecting ink obliterations, while Khan and colleagues explored near-infrared hyperspectral imaging with Principal Component Analysis (PCA) to effectively identify document forgeries.\u003c/p\u003e\u003cp\u003eSubsequent approaches adopted advanced dimensionality reduction techniques. Devassy and George [6] utilized t-distributed stochastic neighbor embedding (t-SNE) to visually cluster spectrally similar inks, aiding manual review. Addressing class imbalance—a common issue when certain inks are underrepresented—Nasir et al. [7] proposed hyperspectral unmixing methods to improve mismatch detection in such challenging conditions.\u003c/p\u003e\u003cp\u003eThe incorporation of deep learning has significantly advanced HSI capabilities. M. J. Khan et al. [8] pioneered the use of Convolutional Neural Networks (CNNs) for ink classification, achieving up to 98% accuracy in blue ink detection. Their subsequent research refined deep learning architectures for more complex scenarios, including black ink differentiation [9], establishing state-of-the-art benchmarks.\u003c/p\u003e\u003cp\u003eTo enhance efficiency without compromising accuracy, Uzair et al. [10] introduced joint sparse band selection techniques. These approaches reduced computational burden by selecting only the most discriminative spectral bands for ink detection, ensuring faster yet reliable processing. Complementary contributions such as those by Khan et al. [5] proposed fuzzy clustering and deep models for forgery detection, while a hybrid spatio-spectral CNN model was also introduced for document authentication [6].\u003c/p\u003e\u003cp\u003eNewer frameworks further enhance accuracy by combining classification and segmentation strategies. For instance, Saleem et al. [11] introduced a patch-based hybrid deep learning model that improves ink-region localization using spectral-spatial cues. Additionally, the move toward real-time applications has begun. Raza et al. [12] developed a portable HSI system integrated with lightweight models for in-field document verification, addressing deployment challenges.\u003c/p\u003e\u003cp\u003eThe field has also benefited from the creation of benchmark datasets, such as the iVision HHID by Ammad ul Islam et al. [8], which supports standardized evaluation of hyperspectral document analysis techniques. Further, broader reviews by Qureshi et al. [9] and Khan et al. [1] summarize current challenges, opportunities, and future directions in HSI applications.\u003c/p\u003e\u003cp\u003eTogether, these twelve studies chronicle the progression of hyperspectral imaging in document forensics—from early spectral sensing strategies to AI-powered, real-time analytical frameworks. The trajectory reflects a significant evolution toward scalable, automated, and non-invasive forensic tools rooted in spectral precision and machine intelligence.\u003c/p\u003e"},{"header":"III. Methodology","content":"\u003cp\u003eA. \u003cem\u003eBand Analysis and Visualization (Part a and b)\u003c/em\u003e:\u003c/p\u003e\u003cp\u003eWe loaded an ENVI-format HSI cube using the Python Spectral library. The cube contains 224 spectral bands ranging approximately from 400 nm to 1000 nm. To analyze band content, we visualized the 1st, 30th, 60th, and last band as grayscale images. Visually, earlier bands are darker and highlight ink absorption. Mid-spectrum bands show greater text-background contrast, while higher bands fade, consistent with ink's low infrared reflectance as can be seen in Fig.\u0026nbsp;1.\u003c/p\u003e\u003cp\u003e\u003cem\u003eB. Spectral Response Analysis (Part c)\u003c/em\u003e:\u003c/p\u003e\u003cp\u003eBinary masks were manually created for two lines of text. For each mask, average spectral reflectance was plotted across all bands. Line 1 and line 2 exhibit distinct spectral profiles, especially in the visible range (400-700nm), indicating different inks as shown in Fig.\u0026nbsp;2.\u003c/p\u003e\u003cp\u003e\u003cem\u003eC. Ink Detection via Clustering (Part d)\u003c/em\u003e:\u003c/p\u003e\u003cp\u003eWe applied K-Means clustering with k = 3 to the spectral pixels from the text masks. This unsupervised method grouped pixels based on spectral similarity.\u003c/p\u003e\u003cp\u003eStrengths: Simple, interpretable, fast.\u003c/p\u003e\u003cp\u003eWeaknesses: Assumes spherical clusters, sensitive to initialization, may underperform with overlapping spectra.\u003c/p\u003e\u003cp\u003e\u003cem\u003eD. Color-Labeling of Inks (Part e)\u003c/em\u003e:\u003c/p\u003e\u003cp\u003eCluster labels were mapped to colors to visualize ink regions. This shows clear separation between text lines with different inks as shown in Fig.\u0026nbsp;3.\u003c/p\u003e\u003cp\u003e\u003cem\u003eE. Dimensionality Reduction using PCA and Reapplying Clustering on PCA Output (Part f and g)\u003c/em\u003e:\u003c/p\u003e\u003cp\u003eThe HSI cube was reshaped and normalized, then PCA was applied to reduce 224 bands to 3 components. Each component was visualized and also used for clustering, the results for which can be observed in Fig.\u0026nbsp;4. K-Means was then reapplied on the PCA-transformed data. The results were comparable to the full-band method.\u003c/p\u003e\u003cp id=\"split\"\u003e\u003cstrong\u003eComparison of Clustering Methods\u003c/strong\u003e:\u003c/p\u003e\n\u003cp\u003e- K-Means (Full Bands): Input Shape N x 224 \u0026mdash; Estimated Inks: \u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e- K-Means (PCA): Input Shape N x 3 \u0026mdash; Estimated Inks:\u003c/p\u003e"},{"header":"IV.\tCONCLUSION","content":"\u003cp\u003eThe approach effectively identifies ink clusters using spectral features. PCA reduces computational load without sacrificing accuracy. K-Means offers a strong baseline method. Using all these methods we can analyze the image and identify the different inks that were used within the document in a non-destructive way.\u003c/p\u003e"},{"header":"V.\tFUTURE WORK","content":"\u003cp\u003eFuture improvements include:\u003c/p\u003e\u003cp\u003e- Applying spectral angle mapper or GMM for better clustering.\u003c/p\u003e\u003cp\u003e- Integrating deep autoencoders for non-linear reduction.\u003c/p\u003e\u003cp\u003e- Testing real-world scanned documents with noise.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eMuhammad Jaleed Khan, Hamid Saeed Khan, Adeel Yousaf, Khurram Khurshid, Asad Abbas, \u0026ldquo;Modern Trends in Hyperspectral Image Analysis: A Review,\u0026rdquo; \u003cem\u003eIEEE Access\u003c/em\u003e, 6(1):14118-14119, 2018.\u003c/li\u003e\n\u003cli\u003eZ. Khan, F. Shafait, A. Mian, \u0026ldquo;Hyperspectral Imaging for Ink Mismatch Detection,\u0026rdquo; \u003cem\u003e12th Int\u0026rsquo;l Conf. on Document Analysis and Recognition (ICDAR)\u003c/em\u003e, 2013, pp. 877\u0026ndash;881.\u003c/li\u003e\n\u003cli\u003eA. Abbas, K. Khurshid, F. Shafait, \u0026ldquo;Towards Automated Ink Mismatch Detection in Hyperspectral Document Images,\u0026rdquo; \u003cem\u003e14th ICDAR\u003c/em\u003e, 2017, pp. 1229\u0026ndash;1236.\u003c/li\u003e\n\u003cli\u003eOstrum, C. J., \u0026ldquo;Use of Tunable Filters in Ink Analysis,\u0026rdquo; \u003cem\u003eForensic Science International\u003c/em\u003e, Vol. 122, 2001, pp. 49\u0026ndash;53.\u003c/li\u003e\n\u003cli\u003eM. J. Khan, A. Yousaf, A. Abbas, K. Khurshid, \u0026ldquo;Deep Learning for Automated Forgery Detection in Hyperspectral Document Images,\u0026rdquo; \u003cem\u003eJournal of Electronic Imaging\u003c/em\u003e, 27(5):053001, 2018.\u003c/li\u003e\n\u003cli\u003eM. J. Khan, K. Khurshid, F. Shafait, \u0026ldquo;A Spatio-Spectral Hybrid Convolutional Architecture for Hyperspectral Document Authentication,\u0026rdquo; \u003cem\u003eICDAR\u003c/em\u003e, 2019, pp. 1097\u0026ndash;1102.\u003c/li\u003e\n\u003cli\u003eNasir, A., Khan, M. J., Shafait, F., Khurshid, K., \u0026ldquo;Hyperspectral Unmixing for Ink Discrimination,\u0026rdquo; \u003cem\u003eDigital Image Computing Techniques and Applications (DICTA)\u003c/em\u003e, 2017.\u003c/li\u003e\n\u003cli\u003eAmmad ul Islam, M. J. Khan, M. Asad, H. A. Khan, K. Khurshid, \u0026ldquo;iVision HHID: Handwritten Hyperspectral Images Dataset for Benchmarking Hyperspectral Imaging-based Document Forensic Analysis,\u0026rdquo; \u003cem\u003eData in Brief\u003c/em\u003e, 2022.\u003c/li\u003e\n\u003cli\u003eRizwan Qureshi, Muhammad Uzair, Khurram Khurshid, Hong Yan, \u0026ldquo;Hyperspectral Document Image Processing: Applications, Challenges and Future Prospects,\u0026rdquo; \u003cem\u003ePattern Recognition\u003c/em\u003e, 90(1):12\u0026ndash;22, 2019.\u003c/li\u003e\n\u003cli\u003eM. Uzair, K. Khurshid, M. J. Khan, \u0026ldquo;Joint Sparse Band Selection for Hyperspectral Ink Analysis,\u0026rdquo; \u003cem\u003eIEEE Transactions on Image Processing\u003c/em\u003e, vol. 29, pp. 4761\u0026ndash;4773, 2020.\u003c/li\u003e\n\u003cli\u003eSaleem, H., Yousaf, A., Khan, H. A., \u0026ldquo;Spectral-Spatial Ink Segmentation Using Hybrid Patch-Based Deep Learning for HSI Documents,\u0026rdquo; \u003cem\u003eForensic Science International: Reports\u003c/em\u003e, 2021.\u003c/li\u003e\n\u003cli\u003eRaza, S., Khan, M. J., Khurshid, K., \u0026ldquo;Design and Deployment of a Portable Real-Time Hyperspectral Imaging System for Document Forensics,\u0026rdquo; \u003cem\u003eSensors\u003c/em\u003e, 2023, 23(2):455.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Hyperspectral Images, Principal Component Analysis, K-Means","lastPublishedDoi":"10.21203/rs.3.rs-7129735/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7129735/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis paper presents an analysis of a hyperspectral image (HSI) of a document with the goal of detecting different inks. The approach includes an analysis of the spectral bands, visual and spectral analysis of the text lines, dimensionality reduction using Principal Component Analysis (PCA), and unsupervised clustering using K-Means to estimate ink categories. The methodology is validated through a comparison of clustering before and after PCA. Visual outputs and spectral plots are provided to display observed trends and ink separation results.\u003c/p\u003e","manuscriptTitle":"Using PCA and K-Means Clustering for Ink Discrimination in Hyperspectral Document Imaging","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-07-18 14:44:41","doi":"10.21203/rs.3.rs-7129735/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4ec99feb-7e91-4096-a8d3-dccb4f315a5c","owner":[],"postedDate":"July 18th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":51560504,"name":"Theoretical Computer Science"}],"tags":[],"updatedAt":"2025-07-18T14:44:41+00:00","versionOfRecord":[],"versionCreatedAt":"2025-07-18 14:44:41","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7129735","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7129735","identity":"rs-7129735","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00