Performance analysis of coalesce private or spill shared (CPOSS) replacement strategy overLRU to handle directory entry evictions from the slice last level cache (LLC) using a scalable coherent sparse directory in Multicore processor | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Performance analysis of coalesce private or spill shared (CPOSS) replacement strategy overLRU to handle directory entry evictions from the slice last level cache (LLC) using a scalable coherent sparse directory in Multicore processor Narottam Sahu, Banchhanidhi Dash, Ahmed Abdulhakim Al-Absi, Prasant Kumar Pattnaik This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4226398/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Ideally, we could solve all memory performance problems by storing everything in high performance caches. However, the associated high-cost forces us to keep only a limited amount of working sets of data in these caches and the replacement of disposable blocks of data with the desired one is a fundamental property to caches. Ultimately, the performance of a cache is heavily influenced by the replacement policy it uses. In this paper, we conduct a comparative analysis of a few replacement policies currently in use from traditional strategies to the ones that attempt to emulate optimal replacement and simulate a few replacement strategies to test the performance of multilevel cache by configuring the cache memory subsystem in the Multi2sim simulator. We have focused on the shared nature of the last Level Cache (LLC) which includes many coherence problems. This is where directories and coherence protocols come into play. However, the blocks suffer when the directory entries are themselves evicted from the directory. In this paper, we look into a cache coherence mechanism that instead of evicting the block from the private caches when the directory entry is evicted, rather than storing the entry in the LLC slice. This guarantees freedom from the Directory Eviction Victim Blocks (DEVB). We analyze the performance of our proposed replacement technique, namely, coalesce private or spill shared (CPOSS) by varying the size and associativity of the sparse directory by providing the workload from Parsec, Splash2 and FFTW benchmarks and we observe that the CPOSS outperforms the existing least recently used (LRU) in slice LLC banks in Multicore processors. Physical sciences/Engineering/Electrical and electronic engineering Physical sciences/Optics and photonics/Optical materials and structures Physical sciences/Optics and photonics/Optical physics Physical sciences/Optics and photonics/Optical techniques Physical sciences/Optics and photonics/Other photonics Cache coherence cache Replacement Sparse Directory. Slice LLC bank Directory Eviction Victim Block Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 1. Introduction The best ache replacement is an important aspect for improving the performance of the slice LLC in a Multicore processor. In multilevel hierarchy to improve the utilization of private cache the directory based approach which keeps track of the live cache block and ensures a reduction in cache invalidation when a block is to be evicted from a slice LLC that is shared by all the cores over a single socket .The sparse Directory maintains the coherence state and the position of a block that is privately used by at least one of the core[ 1 ]. The entry for a cache block is evicted when all private cache blocks are removed from the processor cores. The eviction of a block that is currently used by private cores that enter the sparse directory must perform back invalidation of all private blocks where the sparse directory entry is looked up. The private cache blocks invalidated due to directory entry evictions are known as Directory Eviction Victim Blocks (DEVB) [ 2 ]. The “Fig. 1 ” illustrates the sparse directory eviction victim’s block. Suppose we have three directory entries E1, E2 and E3 tracking blocks Block1, Block2, Block3. Let us consider a quad core CPU. Core 0 is currently using Block 1 and Block3. Core 1 is currently using Block1 and Block2. Core 2 is using Block 2 and Block 4 and Core 3 is using all 4 blocks. When other directory entries E4, E5 and E6 tracking blocks 4, 5 and 6 fight for same position held by E1, E2 and E3 then the directory entries E1, E2 and E3 are evicted from the sparse directory in the LLC. When these entries are evicted from the sparse directory, it sends an invalidate request to each of the blocks the corresponding entry is tracking. As a result, the privately held blocks Block 1, Block 2 and Block 3 are forced to be evicted from cores C0, C1, C2 and C3. These blocks which were forced to be evicted from the private caches are called as DEVBs. 2. Literature Review A Multicore processor with an inclusive last level cache has core cache invalidation due to the directory entry removed from sparse directory. In inclusive LLC we found that it does not involve any directory entry eviction. This is because the LRU victimizes the block before their coalesce or spilled directory entry [ 3 ]. The selection of slice LLC replacement is important to keep track of the impact of sparse directory size and LLC size. Multi2Sim is a simulation framework written in C for heterogeneous computing, which includes models for different types of CPU and GPU architectures. In this framework, guest refers to any characteristic of the program being simulated, while host refers to the properties of the simulator itself. The guest code refers to the instructions of the simulated program, and the host code refers to the instructions executed by multi2sim on the user's machine. Multi2Sim models, configures and implements the memory hierarchy, including caches, main memory, and interconnection networks [ 4 ]. The external networks are referred to as external networks, while the internal networks are referred to as internal networks. When a block is shared by multiple cores, the problem of data consistency and integrity comes into question. It may be such that a cache block shared by multiple cores and held in their own private caches is written and read simultaneously by more than one core. In such case, the data in the same block may be different in different cores and if any other block wants to read the data, then it may not have updated information or may have corrupted data. In order to maintain the data consistency across all cores, coherence protocols are fundamental to the caches. The coherency is maintained by the cache controller. Here we also conduct a comparative study between various cache coherency protocols in use today. Multi2Sim is a computer simulation tool that implements coherence protocol called NMOESI [ 5 ][ 6 ]. A set of all possible state transitions for the NMOESI coherence protocol. The “Figure 2 ” shows the various states and the actions triggered by processing devices, as well as the requests launched internally in the memory hierarchy. The state that requires coherence problem to request block data is the I state. The N state and the S state exhibit similar behaviors, except that when an N block is evicted or has a write request, its data is sent to the next-level cache or a neighboring cache respectively [ 7 ][ 8 ][ 9 ]. From a recent study we observe inclusion policy and Prefetching technique for slice LLC bank deliver good performance only in nonexclusive and suffer from performance degradation in inclusive slice LLC due to the usage of traditional Replacement as the current workload and type of applications varying in the access pattern and their size with respect to multi-thread and parallel applications [ 10 ] 3. Proposed Work The proposed work uses the methodology for handling eviction from the sparse Directory: When an entry E is eliminated from the sparse directory, in order to prevent the invalidation of the blocks E is currently tracking from the private caches of the cores, it is kept in the LLC slice. Generally, entry E is kept in the same set of the block it is tracking. If the set is full, different eviction strategies are used to choose the block that is to be replaced with the directory entry E. Another way is to keep entry E in the LLC by replacing the same block it is tracking. Also, we notice that the block that E is supposed to be tracking is not desired to be kept in the LLC when the state of the LLC is Modified(M)/Exclusive(E) as the block will be provided by the core which is holding it. In this case, we can replace the block with its directory entry. This is known as coalesce Private or Spill Shared (CPOSS) replacement. The figure depicts the operation steps that take place when a directory entry E is evicted from the sparse directory in a socket S. We assume that each socket has its own home socket directory. When an entry E is evicted from the sparse directory it is held until its state returns to a stable state and then the cache controller generates a block W which contains the directory entry E as well as write back directory entry message and send the block W to source socket directory S. If the block is not invalidated or it is invalidated but socket S is the only sharer or owner of the block, then the block is written back to the physical memory. If S is not the only owner or sharer of the block then the block B is read from physical memory, the directory entry E is copied from block W to Block B and kept in appropriate position or the segment that is assigned to S in the block B and the block B written back to the physical memory of S. The working and operational steps have been shown in detail in “Fig. 2 ”. 4. Simulation setup and testing environment and workloads In this work we have use Multi2sim simulation framework to evaluate our proposed replacement policy CPOSS. The configuration used for the purposes of this experiment is given in Table 1 . Various workloads like Parsec, Splash2 and FFTW were provided which are given in detail. Parsec is a benchmark suite used for workload that mimics large-scale commercial workloads of different domains. It has several improvements in the features and has a new workload which also support additional parallelization model which simplified the used of PARSEC [ 11 ]. PARSEC has higher scalability and also covers large number of emerging applications. PARSEC was created for improving existing workloads, increase the application coverage of the suite, simplify the use of PARSEC. The emerging application of PARSEC are video gaming and virtual world. The main feature of PARSEC is its newly added 4 workload, one new benchmark, support for the Intel threading and easy to use of the suite. The sparse directory entries eviction using LRU, from LLC slice eviction will be done using CPOSS and at private L1 and L2 it will be using LRU technique.SPLASH-2 is a recently released suit for parallel applications used in shared address-space multiprocessors[ 11 ].FFT(Fast Fourier Transform) is the most frequently used algorithm for scientific purposes, FFTW(Fastest Fourier Transform of the West)is widely used because of its remarkable performance with comparison to vendor-supplied libraries. FFTW is better than vendor-supplied libraries it is outperformed by FFTE (Fastest Fourier Transform of the East) on large transform sizes [ 12 ].FFTW is able to achieve performance portability by evaluating the speed of various code options on the specific architecture it is running and then dynamically selecting the most suitable one during runtime.[ 13 ] Table 1 System Configurations Testing and Simulator Configuration Environment No. of cores LLC Size L2 Size L1I Size L1D Size 4 32MB 128KB 64KB 64KB 8 64MB 256KB 128KB 128KB 16 128MB 512KB 256KB 256KB 32 512MB 8MB 1MB 1MB Sparse Directory Configuration Sparse Directory Eviction Strategies: LRU LLC Slice replacement: CPOSS Block Size: 64 bytes 5. Result Analysis We analyse the performance of difference benchmarks by varying the sparse directory sizes in 1/2,1/8,1/16, /1/32. The Table 2 shows the speedup obtained in various sparse directory configuration in various benchmarks and “Fig. 3 ” visualizes the speedup for parsec, splash2, FFTW benchmarks with scalable sparse directory entries. Table 2 Performance Impact with Varying Sparse Directory entry Benchmark Speedup 1/2 1/8 1/16 1/32 Parsec 0.8 0.91 0.75 0.6 Splash2 0.6 0.9 0.8 0.7 FFTW 0.5 0.8 0.7 0.6 We observe that the best speedup was obtained for the sparse directory size 1/8. For Parsec benchmark, the respective speedup for sparse directory sizes 1/2,1/8,1/16 and 1/32 was 0.8, 0.91, 0.75 and 0.6. Similar results were obtained for Splash2 benchmarks with speedup being 0.6, 0.9, 0.8 and 0.7. For FFTW benchmark, the speedup was 0.5, 0.8, 0.7 and 0.6. The Table 3 shows speedup for different benchmarks by varying no. of cores and the set associativity of the sparse directory. Table 3 Performance Loss in the reduction of Set-associativity in 4,8,16,32 cores No. of cores LLC Associativity Speedup in various Benchmarks Parsec Splash2 FFTW 4-core 4-way 0.68 0.82 0.78 8-way 0.78 0.87 0.84 16-way 0.82 0.98 0.96 32-way 1.04 1.02 1.03 8-core 4-way 0.72 0.84 0.76 8-way 0.82 0.89 0.86 16-way 0.88 0.99 0.97 32-way 1.24 1.18 1.12 16-core 4-way 0.86 0.82 0.84 8-way 0.92 0.94 0.9 16-way 1.38 1.24 1.16 32-way 1.58 1.46 1.32 32-core 4-way 0.88 0.86 0.84 8-way 0.98 0.92 0.93 16-way 1.04 1.02 1.2 32-way 1.88 1.76 1.98 The best results were obtained for 32-way set associativity in all configurations of no. of cores that were simulated. The maximum speedup was obtained for 32-cores and 32-way set associativity of sparse directory which is evident as the benchmarks have been designed to provide multi-threaded and Multicore workloads. A trend line was found that the speedup increased with increase in associativity of the sparse directory. The figures below show the percentage speedup decrease with respect to 32-way sparse directory in multi- cores. The Fig. 4 ”, “Fig. 5 ”, “Fig. 6 ”, and “Fig. 7 show the percentage speedup decrease in 4-cores, 8-cores, 16-cores and 32-cores respectively. In 4-cores, the speedup decrease was between 25–35% for 4-way, 15–25% for 8-way, 7–22% for 16-way between Parsec, Splash2 and FFTW benchmarks. Similar pattern was found for 8-cores with speedup decrease being 28–48%, 24–34%, and 14–30% for 4-way, 8-way, and 16-way respectively. The same for 16-cores were found to be 37–46%, 32–42%, 13–17% and for 32-cores were found to be 52–58%, 48–54%, and 40–45% for 4-way, 8-way and 16-way sparse directories respectively. The “Fig. 8 ” shows the percentage speed up increase due to increase the capacity of LLC. The speedup of CPOSS relative to LRU replacement for three different benchmarks has been estimated and we observe that with an exclusive LLC size of 512MB parsec benchmark shows better performance as compared to splash2 and FFTW benchmark in CPOSS over LRU. 6. Conclusion The impact of varying the sparse directory sizes and associativity on the speedup was analysed. We observe that we have achieved the best performance when the size of the sparse directory was (1/8) the number of entries with respect to the size of the L2 cache and a general trend was found in which the performance improved with increasing associativity of the sparse directory. In our analysis, the benchmarks performed best in a 32-way sparse directory of size 1/8. We conducted this analysis by using the NMOESI protocol to handle coherency and used our proposed CPOSS replacement in slice LLC,which outperforms the LRU with an exclusive LLC size of 512 MB against the traditional LRU eviction technique In the future we would like to extend our work to compare the performance and efficiency in space utilization of slice LLC with respect to many other cache coherence protocols and state of the art replacement techniques. Declarations Author Contribution All authors contributed equally to this work Data Availability The datasets generated and/or analysed during the current study are available in the [PARSEC benchmark] repository, [https://github.com/Multi2Sim/m2s-bench-parsec-3.0-src]. The datasets generated and/or analysed during the current study are available in the [SPLASH2] repository, [https://github.com/Multi2Sim/m2s-bench-parsec-3.0-src]. The datasets generated and/or analysed during the current study are available in the [FFTW benchmark] repository[https://github.com/Multi2Sim/m2s-bench-splash3]. The data that support the findings of this study are also available from [http://www.multi2sim.org]. The datasets generated and/or analysed during the current study are publicly available. References Chaudhuri, M. Zero directory eviction victim: Unbounded coherence directory and core cache isolation. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) (pp. 277–290). IEEE. (2021). Chaudhuri, M. Zero inclusion victim: isolating core caches from inclusive last-level cache evictions. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) (pp. 71–84). IEEE. (2021,). Qiu, Y., Jiao, J., Zeng, X., & Fan, Y. Tag-Sharer-Fusion Directory: A Scalable Coherence Directory With Flexible Entry Formats. IEEE Transactions on Parallel and Distributed Systems, 34(1), 262–274. (2022). Shukur, H., Zeebaree, S., Zebari, R., Ahmed, O., Haji, L., & Abdulqader, D. Cache coherence protocols in distributed systems. Journal of Applied Science and Technology Trends, 1(3), 92–97(2020). Al-Waisi, Z., & Agyeman, M. O. An overview of on-chip cache coherence protocols”.In2017 intelligent systems conference (IntelliSys) (pp. 304–309). IEEE.Alkhamisi, K. (2022). Cache Coherence issues and Solution: A Review. International Journal of Information Systems and Computer Technologies, 1(2) (2017, September). Nair, A. S., Pai, A. V., Raveendran, B. K., & Patil, G. MOESI L: cache coherency protocol for locked mixed criticality L1 data cache. In 2021 IEEE/ACM 25th International Symposium on Distributed Simulation and Real Time Applications (DS-RT) (pp. 1–8). IEEE (2021, September).. Al-Hothali, S., Soomro, S., Tanvir, K., & Tuli, R. Snoopy anddirectory based cachecoherence protocols: A critical analysis. Journal of Information & Communication Technology, 4(1), 01–10 (2010).. Kayi, A., & El-Ghazawi, T. An adaptive cache coherence protocol for chip multiprocessors. In Proceedings of the Second International Forum on Next-Generation Multicore/Many-core Technologies (pp. 1–10) (2010). Axtmann, M., Witt, S., Ferizovic, D., & Sanders, P. Engineering in-place (shared-memory) sorting algorithms. ACM Transactions on Parallel Computing, 9(1), 1–62 (2022). Sahu, N., Dash, B., Pattnaik, P.K., Bandyopadhyay, A. An Adaptive Replacement Strategy LWIRR for Shared Last Level Cache L3 in Multicore Processors. In: Mahmud, M., Mendoza-Barrera, C., Kaiser, M.S., Bandyopadhyay, A., Ray, K., Lugo, E. (eds) Proceedings of Trends in Electronics and Health Informatics. TEHI 2022. Lecture Notes in Networks and Systems, vol 675.Springer, Singapore. https://doi.org/10.1007/978-981-99-1916-1_31 (2023). Bucek, J., Lange, K. D., & v. Kistowski, J. SPEC CPU2017: Next-generation compute benchmark.In Companion of the 2018 ACM/SPEC International Conference on Performance Engineering (pp. 41–42) (2018). Bienia, C., & Li, K. Parsec 2.0: A new benchmark suite for chip-multiprocessors. In Proceedings of the 5th Annual Workshop on Modeling, Bench-marking and Simulation (Vol. 2011, p. 37) (2009, June). Woo, S. C., Ohara, M., Torrie, E., Singh, J. P., & Gupta, A. The SPLASH-2 programs: Characterization and methodological considerations. ACM SIGARCH computer architecture news, 23(2), 24–36 (1995). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4226398","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":298896156,"identity":"eb877651-d9e5-4de7-99d0-ffc02e4724ed","order_by":0,"name":"Narottam Sahu","email":"","orcid":"","institution":"KIIT University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Narottam","middleName":"","lastName":"Sahu","suffix":""},{"id":298896163,"identity":"5a1a0ba4-0547-420f-9ae9-08d7e686d941","order_by":1,"name":"Banchhanidhi Dash","email":"","orcid":"","institution":"KIIT University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Banchhanidhi","middleName":"","lastName":"Dash","suffix":""},{"id":298896167,"identity":"514c429e-6014-4084-81ca-2590f22471e4","order_by":2,"name":"Ahmed Abdulhakim Al-Absi","email":"","orcid":"","institution":"Kyungdong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ahmed","middleName":"Abdulhakim","lastName":"Al-Absi","suffix":""},{"id":298896170,"identity":"f4cc24b6-06cb-47f8-9628-ede11849833d","order_by":3,"name":"Prasant Kumar Pattnaik","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHklEQVRIie3RsWqDQBjA8U8Cuhy4XrhQX+FEUEItfZUT4bJIaMjS0UmXQNfLW+QRDgpmCWQNZEkQ7JKldBHqUE3TIuVC6dbh/oue5w++UwCd7t9G2dfdDYCVgfxcG6n69UGfMA8AFS1hvxHoE8x7DxQF1rY41A98GuTr49uspo6zrCJ5rMGx00F2UJDxIrbcBU3mo03iEcGou9pz2Q3mCmnkVEGojE2M6GMkIAGCWMMomaQdMVZgZFhFtqU5bDpiv5TviFHmLNdncn+V7GKTIJpEAjOfdAR25nmw6BoZi9InI8rnGJ/8W8Tbs2w4k4zjWDyrSWBH1fDUxFNsT8o9Ctsvlhfeax2Gd095XikHu1zZzw18+WN/IDqdTqf77gM7UGGHEvgLUgAAAABJRU5ErkJggg==","orcid":"","institution":"KIIT University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Prasant","middleName":"Kumar","lastName":"Pattnaik","suffix":""}],"badges":[],"createdAt":"2024-04-06 08:14:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4226398/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4226398/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":55933399,"identity":"b1647ca5-c65d-48ec-a08f-99af54b42fc2","added_by":"auto","created_at":"2024-05-06 13:21:08","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":66085,"visible":true,"origin":"","legend":"\u003cp\u003eSparse Directory Eviction Victim Block in a quad-core Processor\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/3a470c53180117d4b5e35111.png"},{"id":55933398,"identity":"1926aa05-09fd-4581-a62d-fd564b1e6db5","added_by":"auto","created_at":"2024-05-06 13:21:08","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":84971,"visible":true,"origin":"","legend":"\u003cp\u003eOperational Steps in a directory entry eviction from LLC Banks using CPOSS replacement technique.\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/77dc98fc5057e62090ef0374.png"},{"id":55933397,"identity":"20f29459-cda8-47fe-bab9-4fcfca29b8d9","added_by":"auto","created_at":"2024-05-06 13:21:08","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":30737,"visible":true,"origin":"","legend":"\u003cp\u003eSpeedup for parsec,splash2, fftw benchmarks with scalable sparse directory entries\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/8ca6f8322299e6f3a490d5c3.png"},{"id":55933403,"identity":"7a991d72-bb6e-4445-85e4-c0d2728339e6","added_by":"auto","created_at":"2024-05-06 13:21:08","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":19852,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance Loss in the reduction of Set-associativity in 4 cores\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/024c3339961c5bcbdd22c60c.png"},{"id":55933927,"identity":"cfc6c401-32b7-422e-bcf1-a9db979c4b94","added_by":"auto","created_at":"2024-05-06 13:29:08","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":20966,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance Loss in the reduction of Set-associativity in 8 cores\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/ab0f1c266b10f92ae1d2a06f.png"},{"id":55933404,"identity":"75bec8ae-0c17-4ab1-8699-7ddac5e0e901","added_by":"auto","created_at":"2024-05-06 13:21:08","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":22089,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance Loss in the reduction of Set-associativity in 16 cores\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/ffbe389a683fa8818d2bb4ab.png"},{"id":55933400,"identity":"f92cff0e-0965-4dc6-9843-a1b7ae445f33","added_by":"auto","created_at":"2024-05-06 13:21:08","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":20376,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance Loss in the reduction of Set-associativity in 32 cores\u003c/p\u003e","description":"","filename":"image7.png","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/39b19e60d482ad2ced7e8613.png"},{"id":55933928,"identity":"8682652a-133b-426d-a5eb-05f8c3261f4f","added_by":"auto","created_at":"2024-05-06 13:29:08","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":20980,"visible":true,"origin":"","legend":"\u003cp\u003eSpeedup of CPOSS relative to LRU replacement with varying LLC size\u003c/p\u003e","description":"","filename":"image8.png","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/76c7d6ff99add0bec2920ca3.png"},{"id":63049305,"identity":"8aafb090-c258-4ee0-be87-7b0ac0cdb1cf","added_by":"auto","created_at":"2024-08-22 13:36:08","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":700284,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4226398/v1/2d09c08b-4b0d-4363-820d-73a577e79e9c.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Performance analysis of coalesce private or spill shared (CPOSS) replacement strategy overLRU to handle directory entry evictions from the slice last level cache (LLC) using a scalable coherent sparse directory in Multicore processor","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe best ache replacement is an important aspect for improving the performance of the slice LLC in a Multicore processor. In multilevel hierarchy to improve the utilization of private cache the directory based approach which keeps track of the live cache block and ensures a reduction in cache invalidation when a block is to be evicted from a slice LLC that is shared by all the cores over a single socket .The sparse Directory maintains the coherence state and the position of a block that is privately used by at least one of the core[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The entry for a cache block is evicted when all private cache blocks are removed from the processor cores. The eviction of a block that is currently used by private cores that enter the sparse directory must perform back invalidation of all private blocks where the sparse directory entry is looked up. The private cache blocks invalidated due to directory entry evictions are known as Directory Eviction Victim Blocks (DEVB) [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. The \u0026ldquo;Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e\u0026rdquo; illustrates the sparse directory eviction victim\u0026rsquo;s block. Suppose we have three directory entries E1, E2 and E3 tracking blocks Block1, Block2, Block3. Let us consider a quad core CPU. Core 0 is currently using Block 1 and Block3. Core 1 is currently using Block1 and Block2. Core 2 is using Block 2 and Block 4 and Core 3 is using all 4 blocks. When other directory entries E4, E5 and E6 tracking blocks 4, 5 and 6 fight for same position held by E1, E2 and E3 then the directory entries E1, E2 and E3 are evicted from the sparse directory in the LLC. When these entries are evicted from the sparse directory, it sends an invalidate request to each of the blocks the corresponding entry is tracking. As a result, the privately held blocks Block 1, Block 2 and Block 3 are forced to be evicted from cores C0, C1, C2 and C3. These blocks which were forced to be evicted from the private caches are called as DEVBs.\u003c/p\u003e "},{"header":"2. Literature Review","content":"\u003cp\u003eA Multicore processor with an inclusive last level cache has core cache invalidation due to the directory entry removed from sparse directory. In inclusive LLC we found that it does not involve any directory entry eviction. This is because the LRU victimizes the block before their coalesce or spilled directory entry [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. The selection of slice LLC replacement is important to keep track of the impact of sparse directory size and LLC size. Multi2Sim is a simulation framework written in C for heterogeneous computing, which includes models for different types of CPU and GPU architectures. In this framework, guest refers to any characteristic of the program being simulated, while host refers to the properties of the simulator itself. The guest code refers to the instructions of the simulated program, and the host code refers to the instructions executed by multi2sim on the user's machine. Multi2Sim models, configures and implements the memory hierarchy, including caches, main memory, and interconnection networks [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The external networks are referred to as external networks, while the internal networks are referred to as internal networks. When a block is shared by multiple cores, the problem of data consistency and integrity comes into question. It may be such that a cache block shared by multiple cores and held in their own private caches is written and read simultaneously by more than one core. In such case, the data in the same block may be different in different cores and if any other block wants to read the data, then it may not have updated information or may have corrupted data. In order to maintain the data consistency across all cores, coherence protocols are fundamental to the caches. The coherency is maintained by the cache controller. Here we also conduct a comparative study between various cache coherency protocols in use today. Multi2Sim is a computer simulation tool that implements coherence protocol called NMOESI [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e][\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. A set of all possible state transitions for the NMOESI coherence protocol. The \u0026ldquo;Figure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u0026rdquo; shows the various states and the actions triggered by processing devices, as well as the requests launched internally in the memory hierarchy. The state that requires coherence problem to request block data is the I state. The N state and the S state exhibit similar behaviors, except that when an N block is evicted or has a write request, its data is sent to the next-level cache or a neighboring cache respectively [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e][\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e][\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. From a recent study we observe inclusion policy and Prefetching technique for slice LLC bank deliver good performance only in nonexclusive and suffer from performance degradation in inclusive slice LLC due to the usage of traditional Replacement as the current workload and type of applications varying in the access pattern and their size with respect to multi-thread and parallel applications [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/p\u003e"},{"header":"3. Proposed Work","content":"\u003cp\u003eThe proposed work uses the methodology for handling eviction from the sparse Directory: When an entry E is eliminated from the sparse directory, in order to prevent the invalidation of the blocks E is currently tracking from the private caches of the cores, it is kept in the LLC slice. Generally, entry E is kept in the same set of the block it is tracking. If the set is full, different eviction strategies are used to choose the block that is to be replaced with the directory entry E. Another way is to keep entry E in the LLC by replacing the same block it is tracking. Also, we notice that the block that E is supposed to be tracking is not desired to be kept in the LLC when the state of the LLC is Modified(M)/Exclusive(E) as the block will be provided by the core which is holding it. In this case, we can replace the block with its directory entry. This is known as coalesce Private or Spill Shared (CPOSS) replacement. The figure depicts the operation steps that take place when a directory entry E is evicted from the sparse directory in a socket S. We assume that each socket has its own home socket directory. When an entry E is evicted from the sparse directory it is held until its state returns to a stable state and then the cache controller generates a block W which contains the directory entry E as well as write back directory entry message and send the block W to source socket directory S. If the block is not invalidated or it is invalidated but socket S is the only sharer or owner of the block, then the block is written back to the physical memory. If S is not the only owner or sharer of the block then the block B is read from physical memory, the directory entry E is copied from block W to Block B and kept in appropriate position or the segment that is assigned to S in the block B and the block B written back to the physical memory of S. The working and operational steps have been shown in detail in \u0026ldquo;Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u0026rdquo;.\u003c/p\u003e "},{"header":"4. Simulation setup and testing environment and workloads","content":"\u003cp\u003eIn this work we have use Multi2sim simulation framework to evaluate our proposed replacement policy CPOSS. The configuration used for the purposes of this experiment is given in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Various workloads like Parsec, Splash2 and FFTW were provided which are given in detail. Parsec is a benchmark suite used for workload that mimics large-scale commercial workloads of different domains. It has several improvements in the features and has a new workload which also support additional parallelization model which simplified the used of PARSEC [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. PARSEC has higher scalability and also covers large number of emerging applications. PARSEC was created for improving existing workloads, increase the application coverage of the suite, simplify the use of PARSEC. The emerging application of PARSEC are video gaming and virtual world. The main feature of PARSEC is its newly added 4 workload, one new benchmark, support for the Intel threading and easy to use of the suite. The sparse directory entries eviction using LRU, from LLC slice eviction will be done using CPOSS and at private L1 and L2 it will be using LRU technique.SPLASH-2 is a recently released suit for parallel applications used in shared address-space multiprocessors[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].FFT(Fast Fourier Transform) is the most frequently used algorithm for scientific purposes, FFTW(Fastest Fourier Transform of the West)is widely used because of its remarkable performance with comparison to vendor-supplied libraries. FFTW is better than vendor-supplied libraries it is outperformed by FFTE (Fastest Fourier Transform of the East) on large transform sizes [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].FFTW is able to achieve performance portability by evaluating the speed of various code options on the specific architecture it is running and then dynamically selecting the most suitable one during runtime.[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]\u003c/p\u003e\u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSystem Configurations\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003eTesting and Simulator Configuration Environment\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo. of cores\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLLC Size\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eL2 Size\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eL1I Size\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eL1D Size\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e4\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e32MB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e128KB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64KB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e64KB\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e8\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e64MB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e256KB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e128KB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e128KB\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e16\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e128MB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e512KB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e256KB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e256KB\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003e32\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e512MB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8MB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1MB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1MB\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSparse Directory Configuration\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSparse Directory Eviction Strategies: LRU\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eLLC Slice replacement: CPOSS\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eBlock Size: 64 bytes\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e"},{"header":"5. Result Analysis","content":"\u003cp\u003eWe analyse the performance of difference benchmarks by varying the sparse directory sizes in 1/2,1/8,1/16, /1/32. The Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows the speedup obtained in various sparse directory configuration in various benchmarks and \u0026ldquo;Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e\u0026rdquo; visualizes the speedup for parsec, splash2, FFTW benchmarks with scalable sparse directory entries.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance Impact with Varying Sparse Directory entry\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eBenchmark\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c5\" namest=\"c2\"\u003e \u003cp\u003eSpeedup\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1/2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1/8\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1/16\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1/32\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eParsec\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSplash2\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eFFTW\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe observe that the best speedup was obtained for the sparse directory size 1/8. For Parsec benchmark, the respective speedup for sparse directory sizes 1/2,1/8,1/16 and 1/32 was 0.8, 0.91, 0.75 and 0.6. Similar results were obtained for Splash2 benchmarks with speedup being 0.6, 0.9, 0.8 and 0.7. For FFTW benchmark, the speedup was 0.5, 0.8, 0.7 and 0.6. The Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows speedup for different benchmarks by varying no. of cores and the set associativity of the sparse directory.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance Loss in the reduction of Set-associativity in 4,8,16,32 cores\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eNo. of cores\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eLLC Associativity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"3\" nameend=\"c5\" namest=\"c3\"\u003e \u003cp\u003eSpeedup in various Benchmarks\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eParsec\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSplash2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFFTW\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003e\u003cb\u003e4-core\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e4-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e8-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e16-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e32-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.03\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003e\u003cb\u003e8-core\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e4-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.76\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e8-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e16-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e32-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003e\u003cb\u003e16-core\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e4-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e8-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e16-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e32-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.32\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003e\u003cb\u003e32-core\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e4-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.84\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e8-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e16-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e32-way\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.98\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe best results were obtained for 32-way set associativity in all configurations of no. of cores that were simulated. The maximum speedup was obtained for 32-cores and 32-way set associativity of sparse directory which is evident as the benchmarks have been designed to provide multi-threaded and Multicore workloads. A trend line was found that the speedup increased with increase in associativity of the sparse directory. The figures below show the percentage speedup decrease with respect to 32-way sparse directory in multi- cores. The Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e\u0026rdquo;, \u0026ldquo;Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e\u0026rdquo;, \u0026ldquo;Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e\u0026rdquo;, and \u0026ldquo;Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e show the percentage speedup decrease in 4-cores, 8-cores, 16-cores and 32-cores respectively. In 4-cores, the speedup decrease was between 25\u0026ndash;35% for 4-way, 15\u0026ndash;25% for 8-way, 7\u0026ndash;22% for 16-way between Parsec, Splash2 and FFTW benchmarks. Similar pattern was found for 8-cores with speedup decrease being 28\u0026ndash;48%, 24\u0026ndash;34%, and 14\u0026ndash;30% for 4-way, 8-way, and 16-way respectively. The same for 16-cores were found to be 37\u0026ndash;46%, 32\u0026ndash;42%, 13\u0026ndash;17% and for 32-cores were found to be 52\u0026ndash;58%, 48\u0026ndash;54%, and 40\u0026ndash;45% for 4-way, 8-way and 16-way sparse directories respectively. The \u0026ldquo;Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e\u0026rdquo; shows the percentage speed up increase due to increase the capacity of LLC. The speedup of CPOSS relative to LRU replacement for three different benchmarks has been estimated and we observe that with an exclusive LLC size of 512MB parsec benchmark shows better performance as compared to splash2 and FFTW benchmark in CPOSS over LRU.\u003c/p\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003eThe impact of varying the sparse directory sizes and associativity on the speedup was analysed. We observe that we have achieved the best performance when the size of the sparse directory was (1/8) the number of entries with respect to the size of the L2 cache and a general trend was found in which the performance improved with increasing associativity of the sparse directory. In our analysis, the benchmarks performed best in a 32-way sparse directory of size 1/8. We conducted this analysis by using the NMOESI protocol to handle coherency and used our proposed CPOSS replacement in slice LLC,which outperforms the LRU with an exclusive LLC size of 512 MB against the traditional LRU eviction technique In the future we would like to extend our work to compare the performance and efficiency in space utilization of slice LLC with respect to many other cache coherence protocols and state of the art replacement techniques.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAll authors contributed equally to this work\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets generated and/or analysed during the current study are available in the [PARSEC benchmark] repository, [https://github.com/Multi2Sim/m2s-bench-parsec-3.0-src]. The datasets generated and/or analysed during the current study are available in the [SPLASH2] repository, [https://github.com/Multi2Sim/m2s-bench-parsec-3.0-src]. The datasets generated and/or analysed during the current study are available in the [FFTW benchmark] repository[https://github.com/Multi2Sim/m2s-bench-splash3]. The data that support the findings of this study are also available from [http://www.multi2sim.org]. The datasets generated and/or analysed during the current study are publicly available.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eChaudhuri, M. Zero directory eviction victim: Unbounded coherence directory and core cache isolation. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) (pp. 277\u0026ndash;290). IEEE. (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChaudhuri, M. Zero inclusion victim: isolating core caches from inclusive last-level cache evictions. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) (pp. 71\u0026ndash;84). IEEE. (2021,).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQiu, Y., Jiao, J., Zeng, X., \u0026amp; Fan, Y. Tag-Sharer-Fusion Directory: A Scalable Coherence Directory With Flexible Entry Formats. IEEE Transactions on Parallel and Distributed Systems, 34(1), 262\u0026ndash;274. (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShukur, H., Zeebaree, S., Zebari, R., Ahmed, O., Haji, L., \u0026amp; Abdulqader, D. Cache coherence protocols in distributed systems. Journal of Applied Science and Technology Trends, 1(3), 92\u0026ndash;97(2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAl-Waisi, Z., \u0026amp; Agyeman, M. O. An overview of on-chip cache coherence protocols\u0026rdquo;.In2017 intelligent systems conference (IntelliSys) (pp. 304\u0026ndash;309). IEEE.Alkhamisi, K. (2022). Cache Coherence issues and Solution: A Review. International Journal of Information Systems and Computer Technologies, 1(2) (2017, September).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNair, A. S., Pai, A. V., Raveendran, B. K., \u0026amp; Patil, G. MOESI L: cache coherency protocol for locked mixed criticality L1 data cache. In 2021 IEEE/ACM 25th International Symposium on Distributed Simulation and Real Time Applications (DS-RT) (pp. 1\u0026ndash;8). IEEE (2021, September)..\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAl-Hothali, S., Soomro, S., Tanvir, K., \u0026amp; Tuli, R. Snoopy anddirectory based cachecoherence protocols: A critical analysis. Journal of Information \u0026amp; Communication Technology, 4(1), 01\u0026ndash;10 (2010)..\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKayi, A., \u0026amp; El-Ghazawi, T. An adaptive cache coherence protocol for chip multiprocessors. In Proceedings of the Second International Forum on Next-Generation Multicore/Many-core Technologies (pp. 1\u0026ndash;10) (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAxtmann, M., Witt, S., Ferizovic, D., \u0026amp; Sanders, P. Engineering in-place (shared-memory) sorting algorithms. ACM Transactions on Parallel Computing, 9(1), 1\u0026ndash;62 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSahu, N., Dash, B., Pattnaik, P.K., Bandyopadhyay, A. An Adaptive Replacement Strategy LWIRR for Shared Last Level Cache L3 in Multicore Processors. In: Mahmud, M., Mendoza-Barrera, C., Kaiser, M.S., Bandyopadhyay, A., Ray, K., Lugo, E. (eds) Proceedings of Trends in Electronics and Health Informatics. TEHI 2022. Lecture Notes in Networks and Systems, vol 675.Springer, Singapore. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-981-99-1916-1_31\u003c/span\u003e\u003cspan address=\"10.1007/978-981-99-1916-1_31\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBucek, J., Lange, K. D., \u0026amp; v. Kistowski, J. SPEC CPU2017: Next-generation compute benchmark.In Companion of the 2018 ACM/SPEC International Conference on Performance Engineering (pp. 41\u0026ndash;42) (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBienia, C., \u0026amp; Li, K. Parsec 2.0: A new benchmark suite for chip-multiprocessors. In Proceedings of the 5th Annual Workshop on Modeling, Bench-marking and Simulation (Vol. 2011, p. 37) (2009, June).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWoo, S. C., Ohara, M., Torrie, E., Singh, J. P., \u0026amp; Gupta, A. The SPLASH-2 programs: Characterization and methodological considerations. ACM SIGARCH computer architecture news, 23(2), 24\u0026ndash;36 (1995).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Cache coherence, cache Replacement, Sparse Directory. Slice LLC bank, Directory Eviction Victim Block","lastPublishedDoi":"10.21203/rs.3.rs-4226398/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4226398/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIdeally, we could solve all memory performance problems by storing everything in high performance caches. However, the associated high-cost forces us to keep only a limited amount of working sets of data in these caches and the replacement of disposable blocks of data with the desired one is a fundamental property to caches. Ultimately, the performance of a cache is heavily influenced by the replacement policy it uses. In this paper, we conduct a comparative analysis of a few replacement policies currently in use from traditional strategies to the ones that attempt to emulate optimal replacement and simulate a few replacement strategies to test the performance of multilevel cache by configuring the cache memory subsystem in the Multi2sim simulator. We have focused on the shared nature of the last Level Cache (LLC) which includes many coherence problems. This is where directories and coherence protocols come into play. However, the blocks suffer when the directory entries are themselves evicted from the directory. In this paper, we look into a cache coherence mechanism that instead of evicting the block from the private caches when the directory entry is evicted, rather than storing the entry in the LLC slice. This guarantees freedom from the Directory Eviction Victim Blocks (DEVB). We analyze the performance of our proposed replacement technique, namely, coalesce private or spill shared (CPOSS) by varying the size and associativity of the sparse directory by providing the workload from Parsec, Splash2 and FFTW benchmarks and we observe that the CPOSS outperforms the existing least recently used (LRU) in slice LLC banks in Multicore processors.\u003c/p\u003e","manuscriptTitle":"Performance analysis of coalesce private or spill shared (CPOSS) replacement strategy overLRU to handle directory entry evictions from the slice last level cache (LLC) using a scalable coherent sparse directory in Multicore processor","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-05-06 13:21:03","doi":"10.21203/rs.3.rs-4226398/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"3181d980-4cc2-432f-bc67-ee5a3166a8fc","owner":[],"postedDate":"May 6th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":31533170,"name":"Physical sciences/Engineering/Electrical and electronic engineering"},{"id":31533171,"name":"Physical sciences/Optics and photonics/Optical materials and structures"},{"id":31533172,"name":"Physical sciences/Optics and photonics/Optical physics"},{"id":31533174,"name":"Physical sciences/Optics and photonics/Optical techniques"},{"id":31533176,"name":"Physical sciences/Optics and photonics/Other photonics"}],"tags":[],"updatedAt":"2024-08-22T13:28:02+00:00","versionOfRecord":[],"versionCreatedAt":"2024-05-06 13:21:03","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4226398","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4226398","identity":"rs-4226398","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.