Human Latent Metrics: Perceptual and Cognitive Response Corresponds to Distance in GAN Latent Space | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Human Latent Metrics: Perceptual and Cognitive Response Corresponds to Distance in GAN Latent Space Shunichi Kasahara, Naoto Ienaga, Kye Shimizu, Kazuma Takada, Maki Sugimoto This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1339104/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Generative adversarial networks (GANs) generate high-dimensional vector spaces (latent spaces) that can interchangeably represent vectors as images. Advancements have extended their ability to computationally generate images indistinguishable from real images such as faces, and more importantly, manipulate images using their inherit vector values in the latent space. This interchangeability of latent vector has the potential to calculate not only distance in the latent space, but also human perceptual and cognitive distance toward images, i.e., how humans perceive and recognize images. However, it is still unclear how the distance in the latent space corresponds to human perception and cognition. Our studies investigated the correspondence between latent vectors and human perception or cognition through psycho-visual experiments that manipulates the latent vectors of face images. In the perception study, a change perception (CP) task was utilized to examine whether participants could perceive visual changes in face images before and after moving an arbitrary distance in the latent space. In the cognition study, a face cognition (FC) task was utilized to examine whether the participants could recognize a face as the same, even after moving an arbitrary distance in the latent space. The results showed that CP and cognition for face images clearly correlates to the distance in the latent space, which can be modeled with a logistic function. We also investigated how the internal layered structure of the latent space correlates to human response by calculating the regression residual error in each layer. As a result, we observed different residual error trends pertaining to CP and FC. Our experiments show that the distance between face images in the latent space corresponds to human perception and cognition for visual changes in face imagery, and additionally indicates that perception and cognition correspond with the latent space differently. By utilizing our methodology, it will be possible to interchangeably convert between the distance in the latent space and the metric of human perception and cognition, potentially leading to image processing that better reflects human perception and cognition. Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1339104","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":85454613,"identity":"0d06307e-3e9b-48e7-946e-a8fa18b00095","order_by":0,"name":"Shunichi Kasahara","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA90lEQVRIiWNgGAWjYBACCSBmBjH4GRgYDyDEeYjQItnAwECiFoMDKFrwAMn24w8/F9Tckzc+fvzBAYY/NnkGB5gfPmCQuYNTizRPjrH0jGPFhtvO5BgcYGxLKzY4wGZswMDzDKcWOYYcoDa2BMZtN3gYDjA2HE7ccICHTYKB5zBuLfzPH//m+Zdgv3kGO8hhRGiRlkgwk+ZtS0jcIAEKATYitEjOeGNmzduXkDwD5JfEtrTEmYeBfknA4xeJ8+mPb/N8S7DtBwbdgw9/bBL7jjc/fPCxB3eIoYIEIFYAOSmx5wCRWkBAvgFE/iBFyygYBaNgFAxzAAAM6Febomf1bgAAAABJRU5ErkJggg==","orcid":"","institution":"Sony Computer Science Laboratories, Inc","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Shunichi","middleName":"","lastName":"Kasahara","suffix":""},{"id":85454609,"identity":"b7913c62-8389-45ee-bdd3-70f6b751c9d9","order_by":1,"name":"Naoto Ienaga","email":"","orcid":"","institution":"University of Tsukuba","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Naoto","middleName":"","lastName":"Ienaga","suffix":""},{"id":85454610,"identity":"a0805a72-c8ca-4d17-9939-89ff544fbb60","order_by":2,"name":"Kye Shimizu","email":"","orcid":"","institution":"Sony Computer Science Laboratories, Inc","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kye","middleName":"","lastName":"Shimizu","suffix":""},{"id":85454611,"identity":"8735360b-5f16-410f-9452-bdfc38b608fd","order_by":3,"name":"Kazuma Takada","email":"","orcid":"","institution":"Sony Computer Science Laboratories, Inc","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kazuma","middleName":"","lastName":"Takada","suffix":""},{"id":85454612,"identity":"05752082-c089-4f8f-8b67-f4515c8fc1b3","order_by":4,"name":"Maki Sugimoto","email":"","orcid":"","institution":"Keio University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Maki","middleName":"","lastName":"Sugimoto","suffix":""}],"badges":[],"createdAt":"2022-02-08 13:29:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1339104/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1339104/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":18541824,"identity":"968c8c75-0623-4eae-a7f3-60fc0153e431","added_by":"auto","created_at":"2022-02-23 20:25:01","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2622321,"visible":true,"origin":"","legend":"","description":"","filename":"LatentMetricsScientificReports0217.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1339104/v1_covered.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Human Latent Metrics: Perceptual and Cognitive Response Corresponds to Distance in GAN Latent Space","fulltext":[{"header":"Full Text","content":"This preprint is available for \u003ca href='/article/rs-1339104/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e."}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-1339104/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1339104/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Generative adversarial networks (GANs) generate high-dimensional vector spaces (latent spaces) that can interchangeably represent vectors as images. Advancements have extended their ability to computationally generate images indistinguishable from real images such as faces, and more importantly, manipulate images using their inherit vector values in the latent space. This interchangeability of latent vector has the potential to calculate not only distance in the latent space, but also human perceptual and cognitive distance toward images, i.e., how humans perceive and recognize images. However, it is still unclear how the distance in the latent space corresponds to human perception and cognition. Our studies investigated the correspondence between latent vectors and human perception or cognition through psycho-visual experiments that manipulates the latent vectors of face images. In the perception study, a change perception (CP) task was utilized to examine whether participants could perceive visual changes in face images before and after moving an arbitrary distance in the latent space. In the cognition study, a face cognition (FC) task was utilized to examine whether the participants could recognize a face as the same, even after moving an arbitrary distance in the latent space. The results showed that CP and cognition for face images clearly correlates to the distance in the latent space, which can be modeled with a logistic function. We also investigated how the internal layered structure of the latent space correlates to human response by calculating the regression residual error in each layer. As a result, we observed different residual error trends pertaining to CP and FC. Our experiments show that the distance between face images in the latent space corresponds to human perception and cognition for visual changes in face imagery, and additionally indicates that perception and cognition correspond with the latent space differently. By utilizing our methodology, it will be possible to interchangeably convert between the distance in the latent space and the metric of human perception and cognition, potentially leading to image processing that better reflects human perception and cognition.","manuscriptTitle":"Human Latent Metrics: Perceptual and Cognitive Response Corresponds to Distance in GAN Latent Space","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-02-23 20:24:53","doi":"10.21203/rs.3.rs-1339104/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"c9484d95-ebf1-447b-98c1-3ee198e909e1","owner":[],"postedDate":"February 23rd, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2022-05-18T12:44:24+00:00","versionOfRecord":[],"versionCreatedAt":"2022-02-23 20:24:53","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1339104","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1339104","identity":"rs-1339104","version":["v1"]},"buildId":"FbvkV6FR0MCFSLy54lSbu","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.