What Drives House Prices? A Linear Regression Approach to Size, Condition, and Features

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract This research examines the key factors that influence house prices, focusing on how size, condition, and structural features contribute to property valuation. A multivariate analysis using a Linear Regression model was conducted to assess the relationships between crucial features such as square footage, number of bedrooms, bathrooms, floors, and property condition. The analysis revealed that square footage and bathrooms exhibit the strongest positive correlations with house prices (both with correlation values of 0.76), indicating their significant impact on property valuation. In contrast, factors like condition and view demonstrated weaker correlations, suggesting a more limited influence. The Linear Regression model achieved an R-squared value of 0.75, explaining 75% of the variation in house prices based on these features. While the model effectively highlights key price determinants, its limitations in handling non-linear relationships and sensitivity to outliers are noted. This study emphasizes the importance of a nuanced, data-driven approach in understanding house price dynamics, offering valuable insights for buyers, sellers, and industry professionals. Future work could explore advanced predictive models and incorporate additional features to enhance forecasting accuracy.
Full text 108,160 characters · extracted from preprint-html · click to expand
What Drives House Prices? A Linear Regression Approach to Size, Condition, and Features | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article What Drives House Prices? A Linear Regression Approach to Size, Condition, and Features Xiaolin Ju, Vaskar Chakma, Misbahul Amin, Joy Arkhid Chakma This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5743165/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This research examines the key factors that influence house prices, focusing on how size, condition, and structural features contribute to property valuation. A multivariate analysis using a Linear Regression model was conducted to assess the relationships between crucial features such as square footage, number of bedrooms, bathrooms, floors, and property condition. The analysis revealed that square footage and bathrooms exhibit the strongest positive correlations with house prices (both with correlation values of 0.76), indicating their significant impact on property valuation. In contrast, factors like condition and view demonstrated weaker correlations, suggesting a more limited influence. The Linear Regression model achieved an R-squared value of 0.75, explaining 75% of the variation in house prices based on these features. While the model effectively highlights key price determinants, its limitations in handling non-linear relationships and sensitivity to outliers are noted. This study emphasizes the importance of a nuanced, data-driven approach in understanding house price dynamics, offering valuable insights for buyers, sellers, and industry professionals. Future work could explore advanced predictive models and incorporate additional features to enhance forecasting accuracy. Artificial Intelligence and Machine Learning Computer Architecture and Engineering House Price Prediction Linear Regression Multivariate Analysis Property Features Market Valuation Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 1. Introduction Regression learning [ 1 , 2 ] is a powerful statistical method used in machine learning [ 3 , 4 , 5 , 6 ] to model the relationship between a dependent variable and one or more independent variables [ 7 ]. In simpler terms, it allows us to understand how changes in certain features (or variables) affect the value of a particular outcome. In house price prediction, regression models [ 8 ] are used to quantify the relationship between a property’s characteristics, such as its size, condition, location, amenities, and market value [ 9 ]. Regression techniques, particularly linear regression [ 10 ], are essential in understanding and predicting real estate prices. Real estate markets are complex and dynamic, where numerous factors influence the final sale price of a house. As measured by square footage or the number of rooms, size is among the most influential factors. However, additional variables, such as the number of bathrooms, the age of the house, its condition, and even its location, all contribute to its price. Linear regression helps model the relationship between these factors and the target variable (the house price), providing a transparent and interpretable way to understand how each factor affects the final value. For example, the model can reveal that for every extra square foot of living space, the price increases by a certain amount or that a house in excellent condition is likely to command a higher price than one in poor condition [ 11 ]. The primary advantage of using regression learning in house price prediction is its simplicity and interoperability [ 12 ]. Linear regression, in particular, produces an equation that allows us to easily quantify the effect of each feature on the price. The coefficient of each feature indicates how much the price will change for a unit change in that feature, all else being equal. This makes the model not only valuable for prediction but also for gaining insight into which factors are the most important drivers of house prices. For example, by using linear regression, we can identify whether the number of bedrooms or the condition of the house has a stronger impact on price, helping homebuyers and real estate professionals make more informed decisions. However, while linear regression is a valuable tool [ 13 , 14 ], it does have limitations. Real-world data often include nonlinear relationships, where the effect of one feature on the price might change depending on the values of other features. For example, the impact of the size of the house on the price might not be linear, as larger homes tend to be priced in different ranges than smaller homes. Additionally, factors such as location, neighborhood quality, or market trends may interact in complex ways that linear regression cannot easily capture. Outliers [ 15 ]—such as properties that are significantly more expensive than others due to unique features—can also distort the results, leading to less accurate predictions. This study aims to apply linear regression to explore the relationships between key property features and house prices, particularly focusing on factors like size, condition, and other significant attributes. By analyzing the effects of these variables on price, this research seeks to provide a deeper understanding of how various features contribute to property values in the real estate market [16, 17]. While linear regression is a relatively simple technique, this study also considers its limitations and how factors such as outliers and non-linear relationships can impact the model’s accuracy. The findings from this research are expected to offer practical implications for various stakeholders in the housing market. For homeowners and prospective buyers, understanding the key features that influence house prices can guide decisions about buying or selling a property [ 18 , 19 , 20 ]. Real estate professionals can use the insights to refine pricing strategies and better assess market trends. Additionally, the results can help policymakers in urban planning and housing policy, offering a clearer understanding of what makes properties more valuable and how different neighborhoods or housing markets are evolving over time. This paper aims to provide a comprehensive analysis of the role that size, condition, and other key features play in determining house prices. By using regression learning, the research will help demystify the pricing process and contribute to more effective and informed decision making for all parties involved in the housing market. 2. Related Research Several existing works have examined the factors influencing house prices using a variety of models. One widely used method is the hedonic pricing model [ 21 , 22 , 23 , 24 , 25 ], which estimates property values based on individual attributes such as size, location, and amenities. However, this approach often overlooks the interactions between these features, limiting its ability to capture the full complexity of the housing market. For example, Rosen [ 26 ] introduced the concept of the hedonic price function, which decomposes house prices based on individual characteristics but does not fully account for multivariate relationships between variables. More recent research has applied machine learning techniques, such as decision trees [ 27 , 28 , 29 , 30 , 31 ], random forests [ 32 , 33 , 34 , 35 ], and gradient boosting methods [ 36 , 37 , 38 , 39 ], to improve the prediction accuracy of house prices. Studies like those by Li and Zhou (2018) [ 40 ] have shown that these models can handle non-linear relationships and interactions between features, but they often operate as “black boxes [ 41 , 42 , 43 , 44 ]” that lack transparency, making it difficult for real estate professionals to interpret the results. Although these models offer high accuracy, the inability to provide clear explanations behind the predictions limits their practical application in decision-making. Our work seeks to bridge the gap between accuracy and interpretability by utilizing a multivariate linear regression model [ 45 , 46 ], which, while simpler, provides a clear understanding of how factors such as size, condition, and features impact house prices. Unlike previous works that often focus on specific regions, our study uses a publicly available dataset from Kaggle [ 47 , 48 ], enhancing the reproducibility and generalizability of our results. By focusing on the balance between transparency and performance, our model addresses the shortcomings of both traditional and modern approaches. Additionally, we acknowledge the limitations posed by outliers and non-linear relationships, areas often overlooked in previous linear models, thus contributing to a more refined understanding of house price prediction. 3. Methodology This study focuses on predicting house prices using a linear regression model, analyzing how house attributes such as size, condition, and features like the number of bedrooms and bathrooms influence property values. The process begins with data collection, where key features, including square footage and the number of rooms, are extracted from the dataset. Preprocessing steps are crucial for preparing the data: missing values are imputed, categorical variables are encoded, and numerical variables are standardized to ensure consistency across features. Feature selection plays a significant role in this methodology. A correlation matrix [ 49 , 50 ] is used to identify variables that have strong relationships with the target variable—house price. Features like square footage, the number of bedrooms and bathrooms, and the condition of the house are selected for their direct impact on price. Correlation analysis also helps ensure that multicollinearity is minimized, ensuring the model’s predictions are not distorted by highly correlated features [ 51 ]. For model development, linear regression is chosen because of its ability to model linear relationships between dependent and independent variables. The data is split into a training set (80%) and a testing set (20%) to evaluate the model’s performance. Cross-validation is also implemented to further validate the model’s ability to generalize to unseen data. To evaluate the accuracy of the predictions, performance metrics such as R-squared (R²), Mean Absolute Error (MAE), and Root Mean Squared Error (RMSE) [ 52 ] are calculated. These metrics provide a clear understanding of how well the model fits the data and how accurately it can predict house prices [ 53 ]. Finally, the interpretation of the model’s coefficients [ 54 ] reveals the influence of each feature on the predicted house prices. Larger coefficients indicate a stronger impact of those features on the price, giving valuable insights into which factors should be prioritized when evaluating property values. 3.1. Data Import and Initial Exploration The dataset consists of house sale data containing features such as: $$\:y\:=\:price\:\left(target\:variable\right)$$ 1 $$\:{X}_{1}=\:bedrooms,\:{X}_{2}=\:bathrooms,\:{X}_{3}=\:sqft\_living,\:.\:.\:.\:\left(input\:features\right)$$ 2 The dataset was loaded using Python’s pandas\ library for further analysis. A snapshot of the first few rows of the data provided initial insight into its structure: $$\:X=\left\{{X}_{1},\:{X}_{2},{X}_{3},\:\dots\:,{X}_{n}\right\}$$ 3 3.2. Data Cleaning and Preprocessing Handling Missing Data: Let \(\:{X}_{missing}\) be the set of features with missing values. For each feature, missing values were filled by the median \(\:\stackrel{\sim}{X},\) calculated as: $$\:\stackrel{\sim}{{X}_{i}}=median\:\left({X}_{i}\right)$$ 4 which ensures that the central tendency of the data is preserved while handling the missing entries. Feature Selection : Not all features are useful for prediction. A subset of features \(\:\left\{{X}_{1},\:{X}_{2},\:\dots\:\dots\:.,{X}_{k}\right\}\) was chosen based on domain knowledge. Irrelevant features [ 55 ], such as the street address, were removed to simplify the model. Data Type Conversion : The \(\:date\) feature was converted to a numerical format, enabling the analysis of time-related trends. 3.3. Exploratory Data Analysis (EDA) Price Distribution : The distribution of the target variable \(\:y\:\left(house\:price\right)\) was examined using a probability density function (PDF) [ 56 ] and visualized using histograms. The distribution of prices was right-skewed, indicating the presence of high-priced houses that can be considered outliers. The PDF is defined as: $$\:f\left(y\right)=\frac{1}{N}+{\sum\:}_{i=1}^{N}\:\left(\delta\:(y-{y}_{i}\right)$$ 5 where \(\:\varvec{N}\) is the number of samples, and \(\:\delta\:(y-{y}_{i)}\) is the Dirac delta function at \(\:y=\:{y}_{i}\) . Scatter Plots : The relationship between house price ​ \(\:\varvec{y}\) and each feature \(\:{\varvec{X}}_{\varvec{i}}\) was visualized using scatter plots. For example, the relationship between ​ \(\:\varvec{y}\) (price) and \(\:{\varvec{X}}_{3}\) (square footage of living area) can be described as: $$\:y=f\:\left({X}_{3}\right)$$ 6 Correlation Matrix : A correlation matrix was calculated to understand the strength of linear relationships [ 57 ] between the variables \(\:{\varvec{X}}_{\varvec{i}}\) . The correlation between two variables \(\:{\varvec{X}}_{\varvec{i}}\) and \(\:{\varvec{X}}_{\varvec{j}}\) is defined by the Pearson correlation coefficient: $$\:{\rho\:}_{ij}=\:\frac{cov\:({X}_{i},{X}_{j})}{\sigma\:{X}_{i}\sigma\:{X}_{j}}$$ 7 where \(\:\varvec{c}\varvec{o}\varvec{v}\:({\varvec{X}}_{\varvec{i}},{\varvec{X}}_{\varvec{j}})\) is the covariance of \(\:{\varvec{X}}_{\varvec{i}}\) and \(\:{\varvec{X}}_{\varvec{j}}\) and \(\:\varvec{\sigma\:}{\varvec{X}}_{\varvec{i}}\) and \(\:\varvec{\sigma\:}{\varvec{X}}_{\varvec{j}}\) are their standard deviations. The correlation matrix was visualized as heatmap, revealing that features like \(\:\varvec{s}\varvec{q}\varvec{f}\varvec{t}\_\varvec{l}\varvec{i}\varvec{v}\varvec{i}\varvec{n}\varvec{g}\) had a high positive correlation with price. 3.4. Outlier Detection Outliers are extreme values of the target variable \(\:\varvec{y}\) ​ or any feature \(\:{\varvec{X}}_{\varvec{i}}\:\) that deviate significantly from the majority of the data. Mathematically, an outlier can be defined as any point where: $$\:{y}_{i}>\:{Q}_{3}+1.5\times\:IQR\:or\:\:\:\:\:\:\:\:{y}_{i}\:<\:{Q}_{1}-1.5\:\times\:IQR$$ 8 where \(\:{Q}_{1}\:\) and \(\:{Q}_{3}\:\) are the first and third quartiles, and \(\:IQR=\:{Q}_{3}-\:{Q}_{1}\:\) is the interquartile range. Outliers were visually identified using box plots and scatter plots, particularly in relation to the price. 3.5. Model Building and Evaluation A predictive model was developed to estimate house prices based on selected features. The target variable is denoted as \(\:y,\:\) and the input feature vector is \(\:X=\left\{{X}_{1},\:{X}_{2},\:\dots\:\dots\:.,{X}_{k}\right\}\) . 3.5.1. Linear Regression Model A linear regression model was chosen for its simplicity. The model assumes a linear relationship between the target variable y and the input features ​ \(\:{X}_{i}\) , modeled as: $$\:y=\:{\beta\:}_{0}+\:{\beta\:}_{1}{X}_{1}+\:{\beta\:}_{2}{X}_{2}+\dots\:+{\beta\:}_{k}{X}_{k}+\:\in\:$$ 9 where \(\:{\beta\:}_{0}\) is the intercept, \(\:{\{\beta\:}_{0},\:{\beta\:}_{1},\dots\:.,{\beta\:}_{k}\}\:\) were estimated by minimizing the sum of squared residuals: $$\:\begin{array}{c}min\\\:\beta\:\end{array}\:\sum\:_{i=1}^{N}{({y}_{i}-\widehat{{y}_{i\:}})}^{2}$$ 10 where \(\:\widehat{{y}_{i\:}}\) is the predicted price for the i-th house, and \(\:{y}_{i}\) is the actual price. 3.5.2. Model Training The dataset was split into a training set ( \(\:{X}_{train},{y}_{train})\) and a training set ( \(\:{X}_{test},\:{y}_{test})\) , with 80% of the data used for training and 20% for testing. The model was trained using the training set. 4. Model Evaluation The performance of the model was evaluated using Root Mean Square Error (RMSE) and R-squared ( \(\:{R}^{2}\) ) metrics: $$\:RMSE=\:\sqrt{\frac{1}{N}\sum\:_{i=1}^{N}{\left({y}_{i}-\widehat{{y}_{i\:}}\right)}^{2}}$$ 11 $$\:{R}^{2}=1-\:\frac{\sum\:_{i=1}^{N}{\left({y}_{i}-\widehat{{y}_{i\:}}\right)}^{2}}{\sum\:_{i=1}^{N}{\left({y}_{i}-\stackrel{-}{y}\right)}^{2}}$$ 12 where \(\:\stackrel{-}{y}\:\) is the mean of the observed prices. The RMSE provides an estimate of the average deviation of the predicted prices from the actual values, while \(\:{R}^{2}\) represents the proportion of variance in \(\:y\) explained by the model. 5. Results and Discussions The analysis of the house price dataset revealed significant trends and relationships between various features and house prices. Through rigorous data preprocessing, exploratory data analysis (EDA), and the development of a linear regression model, key insights emerged that contribute to understanding the factors influencing housing prices in the market. 5.1. House Price Distribution The initial exploratory analysis highlighted that house prices were predominantly right-skewed, with a concentration of properties priced between $ 200,000 and $ 500,000. The histogram depicting this distribution illustrated that while most transactions occurred in the mid-range, the presence of luxury homes significantly affected the overall average price. This skewness indicates that while affordable housing remains prevalent, high-end properties create a disparity in the perceived average market value. 5.2. Correlation Analysis The correlation matrix identified the square footage of living space and the quality grade of homes as the strongest predictors of house prices. With correlation coefficients of 0.70 and 0.66, respectively, these features demonstrated a clear relationship where larger and higher-quality homes were associated with increased prices. This finding aligns with existing literature, which suggests that buyers prioritize size and quality when evaluating property value. 5.3. Correlation Heatmap The heatmap displays the correlation coefficients ranging from − 1 to + 1, where: A value close to + 1 indicates a strong positive correlation, A value close to -1 indicates a strong negative correlation, A value around 0 suggests no correlation. In this analysis, the heatmap highlights several significant correlations: Square Footage of Living Space (sqft_living): The highest positive correlation coefficient of 0.70 was observed, indicating that as the square footage increases, the price of the house tends to rise significantly. This emphasizes the importance of size in determining property values. Quality Grade With a correlation coefficient of 0.66, this feature also showed a strong positive relationship with house prices. Homes with higher quality grades are associated with higher market values, reflecting buyers' preferences for well-constructed and aesthetically appealing properties. Number of Bathrooms This feature exhibited a moderate positive correlation of 0.52, suggesting that more bathrooms generally contribute to higher prices, although the relationship is not as strong as that seen with square footage and quality grade. 5.4. Comparison of Models The comparison of models [ 58 ] highlights the strengths of Linear Regression in house price prediction. Despite its simplicity, Linear Regression demonstrates competitive performance, offering a balance between accuracy and interpretability. Unlike complex models such as Gradient Boosting or Random Forest, Linear Regression provides clear insights into the relationship between features and target variables, making it more suitable for applications where transparency is crucial. Furthermore, Linear Regression is computationally efficient, requiring less time and resources compared to tree-based models, which are prone to overfitting without careful tuning. While advanced models might slightly improve accuracy, the simplicity, speed, and ease of implementation of Linear Regression make it a reliable and practical choice for real-world applications, particularly when interpretability and efficiency are prioritized. Table 1 Comparison of Models. Model MAE MSE R 2 Linear Regression 210,908.173 9.869 × 10 11 0.032284 Decision Tree 262,910.016 1.052 × 10 12 -0.031647 Random Forest 208,109.707 9.917 × 10 11 0.027524 Gradient Boosting 202,521.382 9.814 × 10 11 0.037695 5.5. Implications of Findings The results of this study have important implications for both homebuyers and real estate professionals. Based on the regression coefficients, the following points can be highlighted: Bedrooms: The negative coefficient for bedrooms (-66999.17) suggests that, all else being equal, an increase in the number of bedrooms tends to decrease the house price. This could indicate that buyers prioritize other factors, such as living space, over the number of bedrooms. Table 2 Linear Regression Coefficients. Feature Coefficient bedrooms -66999.169310 bathrooms -7519.141800 sqft_living 311.015539 sqft_lot -0.598012 floors 30932.522099 condition 61097.200192 Bathrooms: The negative coefficient for bathrooms (7519.14) implies that additional bathrooms may not significantly increase the house price, or that buyers may perceive diminishing returns on extra bathroom features. Square Footage of Living Space ( sqft living ): The positive coefficient for sqft living (311.02) indicates that larger living spaces contribute positively to house prices. This suggests that homebuyers value more living space, and homes with larger square footage tend to command higher prices. Lot Size ( sqft lot ): The negative coefficient for sqft lot (0.60) implies that a larger lot size may not have a strong impact on house prices, or that buyers are more focused on the house’s features rather than the size of the lot. Floors: The positive coefficient for floors (30932.52) suggests that houses with more floors are generally valued higher, indicating that multi-story homes may appeal more to buyers, especially if they provide more living space. Condition: The positive coefficient for condition (61097.20) highlights that the condition of a home plays a significant role in determining its price. Homes in better condition tend to fetch higher prices, making it important for sellers to maintain the quality of their properties. These findings provide actionable insights for different stakeholders in the housing market: Homebuyers: Understanding these factors can help buyers prioritize their preferences. For instance, they may place greater importance on the size of the living space and the condition of the house, while less importance on the number of bedrooms and bathrooms. Real Estate Professionals: Agents can use these insights to adjust their marketing strategies, emphasizing properties with larger living spaces, better condition, and more floors as premium features. Policymakers: These results can guide policymakers in addressing housing affordability by focusing on mid-range housing options with better condition and reasonable sizes, which are likely to be in higher demand. 6. Results and Discussions This analysis has highlighted the key factors influencing house prices, including the size of the living space and the quality of the home. Larger homes and those in better condition tend to be priced higher. The number of floors and the overall condition of the property also play significant roles in determining price, with better-maintained homes and those with more floors generally fetching higher values. However, location was not fully addressed in this study, although it is clear that it plays a crucial role in pricing. The presence of high-value outliers in certain neighborhoods suggests that the area where a home is located can significantly impact its price. While the linear regression model worked well in capturing general trends, it may not be the best fit for predicting prices of luxury properties, which require more complex modeling. In the future, expanding the dataset to include economic factors and location-specific details, such as neighborhood amenities, could improve predictions. Advanced modeling techniques like decision trees or machine learning algorithms would also help capture more complex pricing patterns, particularly for luxury homes. This study has some limitations. The linear regression model, while effective, may not fully capture the complexities of luxury properties, which often have unique pricing patterns. Additionally, the dataset used is from one region, so future research could apply this model to different areas to see if the findings hold elsewhere. The study also didn’t consider factors like economic conditions or market changes over time. Including variables such as interest rates, inflation, or local economic data would provide a more complete understanding of house price dynamics. Future research should also explore more advanced techniques, like machine learning, to improve prediction accuracy and account for the unique features of high-end properties. Expanding the dataset and including more variables will help provide a deeper insight into the housing market and improve decision making for buyers, sellers, and policymakers. Declarations Acknowledgment We would like to thank our supervisor for his efforts and guidance. Declaration of Competing Interests The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. CRediT Authorship Contribution Statement Xiaolin Ju: Data curation, Software, Validation, Conceptualization, Methodology, Writing -review & editing, Investigation, Supervision. Vaskar Chakma: Software, Conceptualization, Methodology, Writing -review & editing. Misbahul Amin: Software, Validation, Review & editing. Joy Arkhid Chakma: Validation, Review & editing. References Saunders C, Gammerman A, Vovk V (1998) Ridge regression learning algorithm in dual variables Forys I (2022) Machine learning in house price analysis: regression´ models versus neural networks. Procedia Comput Sci 207:435–445 Jordan MI, Mitchell TM (2015) Machine learning: Trends, perspectives, and prospects. Science 349(6245):255–260 Park B, Bae JK (2015) Using machine learning algorithms for housing price prediction: The case of fairfax county, virginia housing data. Expert Syst Appl 42(6):2928–2934 Truong Q, Nguyen M, Dang H, Mei B (2020) Housing price prediction via improved machine learning techniques. Procedia Comput Sci 174:433–442 Nikou M, Mansourfar G, Bagherzadeh J (2019) Stock price prediction using deep learning algorithm and its comparison with machine learning algorithms. Intell Syst Acc Finance Manage 26(4):164–174 Wang Y-X, Hebert M (2016) Learning to learn: Model regression networks for easy small sample learning, in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part VI 14, pp. 616–634, Springer Gong J, Sun S (2009) A new approach of stock price prediction based on logistic regression model, in 2009 International Conference on New Trends in Information and Service Science, pp. 1366–1371, IEEE Taylor LO, Smith VK (2000) Environmental amenities as a source of market power. Land Econ, pp. 550–568 Maulud D, Abdulazeez AM (2020) A review on linear regression comprehensive in machine learning. J Appl Sci Technol Trends 1(2):140–147 Himmelberg C, Mayer C, Sinai T (2005) Assessing high house prices: Bubbles, fundamentals and misperceptions. J Economic Perspect 19(4):67–92 Li M (2015) Moving beyond the linear regression model: Advantages of the quantile regression model. J Manag 41(1):71–98 Nimon KF, Oswald FL (2013) Understanding the results of multiple linear regression: Beyond standardized regression coefficients. Organizational Res Methods 16(4):650–674 Hocking RR (2013) Methods and applications of linear models: regression and the analysis of variance. Wiley Ghosh D, Vogt A, Outliers: An evaluation of methodologies, in Joint statistical meetings, vol. 12, pp. 3455–3460, [16], Liu C, Xiong W (2012) China’s real estate market, 2018 Kahr J, Thomsett MC (2006) Real estate market valuation and analysis. Wiley Hassan MM, Ahmad N, Hashim AH (2021) The conceptual framework of housing purchase decision-making process. Int J Acad Res Bus Social Sci 11(11):1673–1690 Ullah F, Sepasgozar SM (2020) Key factors influencing purchase or rent decisions in smart real estate investments: A system dynamics approach using online forum thread data. Sustainability 12(11):4382 Duca JV, Muellbauer J, Murphy A (2021) What drives house price cycles? international experience and policy issues. J Econ Lit 59(3):773–864 Chau KW, Chin T (2003) A critical review of literature on the hedonic price model. Int J Hous Sci Appl 27(2):145–165 Nguyen M-LT (2020) The hedonic pricing model applied to the housing market. Int J Econ Bus Adm 8(3):416–428 Zaki J, Nayyar A, Dalal S, Ali ZH (2022) House price prediction using hedonic pricing model and machine learning techniques. Concurrency computation: Pract experience 34(27):e7342 Lisi G (2019) Property valuation: the hedonic pricing model–location and housing submarkets. J Property Invest Finance 37(6):589–596 Heyman AV, Law S, Berghauser Pont M (2018) How is location measured in housing valuation? a systematic review of accessibility specifications in hedonic price models. Urban Sci 3(1):3 Rosen S (1974) Hedonic prices and implicit markets: product differentiation in pure competition. J Polit Econ 82(1):34–55 Rokach L, Maimon O (2005) Decision trees, Data mining and knowledge discovery handbook, pp. 165–192 Zhang Z (2021) Decision trees for objective house price prediction, in 2021 3rd International Conference on Machine Learning, Big Data and Business Intelligence (MLBDBI), pp. 280–283, IEEE Reddy PSM et al (2023) Decision tree regressor compared with random forest regressor for house price prediction in mumbai. J Surv Fisheries Sci 10(1):2323–2332 Thamarai M, Malarvizhi S (2020) House price prediction modeling using machine learning. Int J Inform Eng Electron Bus, 12, 2 Rana VS, Mondal J, Sharma A, Kashyap I (2020) House price prediction using optimal regression techniques, in 2020 2nd International Conference on Advances in Computing, Communication Control and Networking (ICACCCN), pp. 203–208, IEEE Breiman L (2001) Random forests. Mach Learn 45:5–32 Hong J, Choi H, Kim W-s (2020) A house price valuation based on the random forest approach: the mass appraisal of residential property in south korea. Int J Strategic Property Manage 24(3):140–152 Zhang Y, Huang J, Zhang J, Liu S, Shorman S (2022) Analysis and prediction of second-hand house price based on random forest. 7(1):27–42 Applied Mathematics and Nonlinear Sciences Jamil SS, Bansal S, Vinjamuri L (2023) House price prediction using random forest techniques: a comparative study Bentejac C, A. Cs´ org¨ o, and, Mart˝ ´ınez-Munoz G (2021) A comparative˜ analysis of gradient boosting algorithms, Artificial Intelligence Review, vol. 54, pp. 1937–1967 Sibindi R, Mwangi RW, Waititu AG (2023) A boosting ensemble learning based hybrid light gradient boosting machine and extreme gradient boosting model for predicting house prices. Eng Rep 5(4):e12599 Hjort A, Pensar J, Scheel I, Sommervoll DE (2022) House price prediction with gradient boosted trees under different loss functions. J Property Res 39(4):338–364 Li S, Jiang Y, Ke S, Nie K, Wu C (2021) Understanding the effects of influential factors on housing prices by combining extreme gradient boosting and a hedonic price model (xgboosthpm). Land 10(5):533 Wang S, Li H, Li J, Zhang Y, Zou B (2018) Automatic analysis of lateral cephalograms based on multiresolution decision tree regression voting, Journal of healthcare engineering, vol. no. 1, p. 1797502, 2018 Seaman JA (2008) Black boxes. Emory LJ 58:427 Khan M, Debnath P, Al Sayeed A, Sumon MFI, Rahman A, Khan M, Pant L (2024) Explainable ai and machine learning model for california house price predictions: Intelligent model for homebuyers and policymakers. J Bus Manage Stud 6(5):73–84 Beimer J, Francke M et al (2019) Out-of-sample house price prediction by hedonic price models and machine learning algorithms. Real Estate Res Q 18(2):13–20 Mueller-Kett C (2024) Artificial intelligence for greater transparency in housing price estimation. AGILE: GIScience Ser 5:41 Manjula R, Jain S, Srivastava S, Kher PR (2017) Real estate value prediction using multivariate regression models, in IOP Conference Series: Materials Science and Engineering, vol. 263, p. 042098, IOP Publishing Mao Y, Yao R (2020) A geographic feature integrated multivariate linear regression method for house price prediction, in 2020 3rd international conference on humanities education and social sciences (ICHESS 2020), pp. 347–351, Atlantis Press Lu S, Li Z, Qin Z, Yang X, Goh RSM (2017) A hybrid regression technique for house prices prediction, in IEEE international conference on industrial engineering and engineering management (IEEM), pp. 319–323, IEEE, 2017 Yu J (2022) Prediction on housing price based on the data on kaggle, in 2022 3rd International Conference on E-commerce and Internet Technology (ECIT 2022), pp. 627–634, Atlantis Press Chen N (2022) House price prediction model of zhaoqing city based on correlation analysis and multiple linear regression analysis, Wireless Communications and Mobile Computing, vol. no. 1, p. 9590704, 2022 Yıldız S (2023) and P. Kara-¨ dayı Atas¸, A novel hybrid house price prediction model, Computational economics, vol. 62, no. 3, pp. 1215–1232 Yahya N, Zainuddin NMM, Sjarif NNA, Azmi NFM (2020) Correlation analysis of factors affecting the prediction of price of terrace houses in penang, malaysia: A case study. Open Int J Inf 8(2):18–39 Cabuk KS, Cengiz SK, Guler MG, Ozturk H, Efe AC, Ulas MG, Karademir FP (2023) Chasing the objective upper eyelid symmetry formula; r2, rmse, poc, mae and mse Chicco D, Warrens MJ, Jurman G (2021) The coefficient of determination r-squared is more informative than smape, mae, mape, mse and rmse in regression analysis evaluation. Peerj Comput Sci 7:e623 Malpezzi S (1999) A simple error correction model of house prices. J Hous Econ 8(1):27–62 Batory D (2005) Feature models, grammars, and propositional formulas, in International Conference on Software Product Lines, pp. 7–20, Springer Afrasiabi M, Mohammadi M, Rastegar M, Stankovic L, Afrasiabi S, Khazaei M (2020) Deep-based conditional probability density function forecasting of residential loads. IEEE Trans Smart Grid 11(4):3646–3657 Miles J, Shevlin M (2000) Applying regression and correlation: A guide for students and researchers Soegianto LM, Hinandra AT, Suri PA, Fajar M (2024) Comparison of model performance on housing business using linear regression, random forest regressor, svr, and neural network. Procedia Comput Sci 245:1139–1145 Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5743165","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":396255226,"identity":"6ba24d99-c555-415d-bbe7-05d7330372b6","order_by":0,"name":"Xiaolin Ju","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAsklEQVRIiWNgGAWjYBACxgYg8aACwpEgXkvCGYhq4rSAQWIbKVqYZySwSSTOu1NncID54G0eBrs8wg6bkcBskLjtmYTBAbZkax6G5GJitDA+SNx2GKiFx0yah+FAYgMRWoDK5oC08H8jWgvQlgawLWxEaul52GyQcOyw5MzDbMaWcwySCWsxbE8+JvGh5jA/3/HmhzfeVNgRoaWBEaqGGUQYEFIPBPJEqBkFo2AUjIKRDgCTxjkaOIbnggAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-2579-5359","institution":"Nantong University","correspondingAuthor":true,"prefix":"","firstName":"Xiaolin","middleName":"","lastName":"Ju","suffix":""},{"id":396255570,"identity":"07478c5c-69d0-46ea-9b9f-1e1f99d325c8","order_by":1,"name":"Vaskar Chakma","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABIklEQVRIie3NsUrDQBjA8QsHdblya0Kkz3CQRbHUB3FJCNglJ4ovcFO6JLhmal8hIVAcHC5k6GBC1o5CoC4ZBJcGb/AirUss1c3h/nDHfdz9OABUqv8Z7DYbILm/fc2Q7W7s40SLAJGzxnavf0Eg6gg4QshqVZP2aXKDg2fnfSzEFZ5pDN5tc4BPPALaxz4pPMsJNu69XoapSX1Co1ySyM6BETREC4seMZgHOeLcYdUwNikjlHUESULWHoGa3ycPr3UmJFlUw/TjTBC62JPLAwTrtuV2v8RluDTBgND4+xf9EGks65S7TlKUy/PQt2giSY6up0gvNrdZ2CcDPK2Nhk+ceUHT9VaM6LzKsxqNL0Z45iYvbZ/8HJcL7Q8qlUql+nufShxw/pnc/wcAAAAASUVORK5CYII=","orcid":"https://orcid.org/0009-0003-3039-3175","institution":"Nantong University","correspondingAuthor":true,"prefix":"","firstName":"Vaskar","middleName":"","lastName":"Chakma","suffix":""},{"id":396255571,"identity":"d78440e6-d29b-49cd-b595-d8b78beb1b8c","order_by":2,"name":"Misbahul Amin","email":"","orcid":"https://orcid.org/0009-0009-3382-9142","institution":"Nantong University","correspondingAuthor":false,"prefix":"","firstName":"Misbahul","middleName":"","lastName":"Amin","suffix":""},{"id":396255572,"identity":"8b41b1c7-a533-49b3-8ad6-3faef66854b1","order_by":3,"name":"Joy Arkhid Chakma","email":"","orcid":"","institution":"Nagaoka University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Joy","middleName":"Arkhid","lastName":"Chakma","suffix":""}],"badges":[],"createdAt":"2024-12-31 17:05:12","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true},"doi":"10.21203/rs.3.rs-5743165/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5743165/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":72863400,"identity":"92c729ee-1632-4613-ab31-af11390fe57e","added_by":"auto","created_at":"2025-01-03 04:27:48","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":15334,"visible":true,"origin":"","legend":"\u003cp\u003eEnd-to-End Workflow for House Price Analysis and Prediction.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/fa9446920cd21a793070492c.png"},{"id":72862859,"identity":"21feb441-990f-4a33-8f6d-2d769315ba54","added_by":"auto","created_at":"2025-01-03 04:19:48","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":20435,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of House Prices.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/95f6e77100e5f6382a22edee.png"},{"id":72863614,"identity":"a6c3cec0-d64e-43d5-8780-cbeb539c0529","added_by":"auto","created_at":"2025-01-03 04:35:48","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":56947,"visible":true,"origin":"","legend":"\u003cp\u003eRelationship Between Sqft living and Price\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/e48ace81e3bfdad76530778e.png"},{"id":72862873,"identity":"c085b418-d43e-4e3c-8f35-1ea561abda59","added_by":"auto","created_at":"2025-01-03 04:19:49","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":34213,"visible":true,"origin":"","legend":"\u003cp\u003eRelationship Between Bedrooms and Price\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/e4e82d281bd52b6102fb2f61.png"},{"id":72863615,"identity":"9980a596-c015-462d-8318-1c6caceeacb8","added_by":"auto","created_at":"2025-01-03 04:35:49","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":22451,"visible":true,"origin":"","legend":"\u003cp\u003eRelationship Between Floors and Price.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/2f346742f17c6df77879027e.png"},{"id":72863403,"identity":"1db35740-7291-4c4a-a7d2-7fbf91e2922b","added_by":"auto","created_at":"2025-01-03 04:27:48","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":47352,"visible":true,"origin":"","legend":"\u003cp\u003eRelationship Between Bathrooms and Price.\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/c161c0607112668d48e07bba.png"},{"id":72863402,"identity":"9a24e8a6-a103-4a90-ba74-4fbe93ecebc9","added_by":"auto","created_at":"2025-01-03 04:27:48","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":20777,"visible":true,"origin":"","legend":"\u003cp\u003eRelationship Between Condition and Price.\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/a44ef78239b3620ccd5eb10f.png"},{"id":72862876,"identity":"3028f1be-a71b-4cc0-b967-cb13b4c0f941","added_by":"auto","created_at":"2025-01-03 04:19:49","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":76352,"visible":true,"origin":"","legend":"\u003cp\u003eCorrelation Heatmap.\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/fbb6cc88354616492b9e931d.png"},{"id":72863404,"identity":"34d89ad9-00a8-46b3-a65c-18f098d48548","added_by":"auto","created_at":"2025-01-03 04:27:48","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":7343,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of Models.\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/8e47b57535cc23d5272e659e.png"},{"id":72864151,"identity":"8dd17701-83aa-45fc-88e9-cc94dd23b068","added_by":"auto","created_at":"2025-01-03 04:43:50","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":876257,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5743165/v1/27fd80d1-719b-4d59-8ce3-4100544cbddd.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eWhat Drives House Prices? A Linear Regression Approach to Size, Condition, and Features\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eRegression learning [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] is a powerful statistical method used in machine learning [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] to model the relationship between a dependent variable and one or more independent variables [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. In simpler terms, it allows us to understand how changes in certain features (or variables) affect the value of a particular outcome. In house price prediction, regression models [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] are used to quantify the relationship between a property\u0026rsquo;s characteristics, such as its size, condition, location, amenities, and market value [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eRegression techniques, particularly linear regression [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], are essential in understanding and predicting real estate prices. Real estate markets are complex and dynamic, where numerous factors influence the final sale price of a house. As measured by square footage or the number of rooms, size is among the most influential factors. However, additional variables, such as the number of bathrooms, the age of the house, its condition, and even its location, all contribute to its price. Linear regression helps model the relationship between these factors and the target variable (the house price), providing a transparent and interpretable way to understand how each factor affects the final value. For example, the model can reveal that for every extra square foot of living space, the price increases by a certain amount or that a house in excellent condition is likely to command a higher price than one in poor condition [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe primary advantage of using regression learning in house price prediction is its simplicity and interoperability [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Linear regression, in particular, produces an equation that allows us to easily quantify the effect of each feature on the price. The coefficient of each feature indicates how much the price will change for a unit change in that feature, all else being equal. This makes the model not only valuable for prediction but also for gaining insight into which factors are the most important drivers of house prices. For example, by using linear regression, we can identify whether the number of bedrooms or the condition of the house has a stronger impact on price, helping homebuyers and real estate professionals make more informed decisions. However, while linear regression is a valuable tool [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e], it does have limitations. Real-world data often include nonlinear relationships, where the effect of one feature on the price might change depending on the values of other features. For example, the impact of the size of the house on the price might not be linear, as larger homes tend to be priced in different ranges than smaller homes. Additionally, factors such as location, neighborhood quality, or market trends may interact in complex ways that linear regression cannot easily capture. Outliers [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]\u0026mdash;such as properties that are significantly more expensive than others due to unique features\u0026mdash;can also distort the results, leading to less accurate predictions.\u003c/p\u003e \u003cp\u003eThis study aims to apply linear regression to explore the relationships between key property features and house prices, particularly focusing on factors like size, condition, and other significant attributes. By analyzing the effects of these variables on price, this research seeks to provide a deeper understanding of how various features contribute to property values in the real estate market [16, 17]. While linear regression is a relatively simple technique, this study also considers its limitations and how factors such as outliers and non-linear relationships can impact the model\u0026rsquo;s accuracy. The findings from this research are expected to offer practical implications for various stakeholders in the housing market. For homeowners and prospective buyers, understanding the key features that influence house prices can guide decisions about buying or selling a property [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Real estate professionals can use the insights to refine pricing strategies and better assess market trends. Additionally, the results can help policymakers in urban planning and housing policy, offering a clearer understanding of what makes properties more valuable and how different neighborhoods or housing markets are evolving over time. This paper aims to provide a comprehensive analysis of the role that size, condition, and other key features play in determining house prices. By using regression learning, the research will help demystify the pricing process and contribute to more effective and informed decision making for all parties involved in the housing market.\u003c/p\u003e"},{"header":"2. Related Research","content":"\u003cp\u003eSeveral existing works have examined the factors influencing house prices using a variety of models. One widely used method is the hedonic pricing model [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e25\u003c/span\u003e], which estimates property values based on individual attributes such as size, location, and amenities. However, this approach often overlooks the interactions between these features, limiting its ability to capture the full complexity of the housing market. For example, Rosen [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e26\u003c/span\u003e] introduced the concept of the hedonic price function, which decomposes house prices based on individual characteristics but does not fully account for multivariate relationships between variables.\u003c/p\u003e \u003cp\u003eMore recent research has applied machine learning techniques, such as decision trees [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e31\u003c/span\u003e], random forests [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e33\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e34\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e35\u003c/span\u003e], and gradient boosting methods [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e37\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e38\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e39\u003c/span\u003e], to improve the prediction accuracy of house prices. Studies like those by Li and Zhou (2018) [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e40\u003c/span\u003e] have shown that these models can handle non-linear relationships and interactions between features, but they often operate as \u0026ldquo;black boxes [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e41\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e42\u003c/span\u003e, \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e43\u003c/span\u003e, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e44\u003c/span\u003e]\u0026rdquo; that lack transparency, making it difficult for real estate professionals to interpret the results. Although these models offer high accuracy, the inability to provide clear explanations behind the predictions limits their practical application in decision-making. Our work seeks to bridge the gap between accuracy and interpretability by utilizing a multivariate linear regression model [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e45\u003c/span\u003e, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e46\u003c/span\u003e], which, while simpler, provides a clear understanding of how factors such as size, condition, and features impact house prices. Unlike previous works that often focus on specific regions, our study uses a publicly available dataset from Kaggle [\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e47\u003c/span\u003e, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e48\u003c/span\u003e], enhancing the reproducibility and generalizability of our results. By focusing on the balance between transparency and performance, our model addresses the shortcomings of both traditional and modern approaches. Additionally, we acknowledge the limitations posed by outliers and non-linear relationships, areas often overlooked in previous linear models, thus contributing to a more refined understanding of house price prediction.\u003c/p\u003e"},{"header":"3. Methodology","content":"\u003cp\u003eThis study focuses on predicting house prices using a linear regression model, analyzing how house attributes such as size, condition, and features like the number of bedrooms and bathrooms influence property values. The process begins with data collection, where key features, including square footage and the number of rooms, are extracted from the dataset. Preprocessing steps are crucial for preparing the data: missing values are imputed, categorical variables are encoded, and numerical variables are standardized to ensure consistency across features.\u003c/p\u003e \u003cp\u003eFeature selection plays a significant role in this methodology. A correlation matrix [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e49\u003c/span\u003e, \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e50\u003c/span\u003e] is used to identify variables that have strong relationships with the target variable\u0026mdash;house price. Features like square footage, the number of bedrooms and bathrooms, and the condition of the house are selected for their direct impact on price. Correlation analysis also helps ensure that multicollinearity is minimized, ensuring the model\u0026rsquo;s predictions are not distorted by highly correlated features [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e51\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFor model development, linear regression is chosen because of its ability to model linear relationships between dependent and independent variables. The data is split into a training set (80%) and a testing set (20%) to evaluate the model\u0026rsquo;s performance. Cross-validation is also implemented to further validate the model\u0026rsquo;s ability to generalize to unseen data. To evaluate the accuracy of the predictions, performance metrics such as R-squared (R\u0026sup2;), Mean Absolute Error (MAE), and Root Mean Squared Error (RMSE) [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e52\u003c/span\u003e] are calculated. These metrics provide a clear understanding of how well the model fits the data and how accurately it can predict house prices [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e53\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFinally, the interpretation of the model\u0026rsquo;s coefficients [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e54\u003c/span\u003e] reveals the influence of each feature on the predicted house prices. Larger coefficients indicate a stronger impact of those features on the price, giving valuable insights into which factors should be prioritized when evaluating property values.\u003c/p\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1. Data Import and Initial Exploration\u003c/h2\u003e \u003cp\u003eThe dataset consists of house sale data containing features such as:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:y\\:=\\:price\\:\\left(target\\:variable\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\:{X}_{1}=\\:bedrooms,\\:{X}_{2}=\\:bathrooms,\\:{X}_{3}=\\:sqft\\_living,\\:.\\:.\\:.\\:\\left(input\\:features\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe dataset was loaded using Python\u0026rsquo;s pandas\\ library for further analysis. A snapshot of the first few rows of the data provided initial insight into its structure:\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$\\:X=\\left\\{{X}_{1},\\:{X}_{2},{X}_{3},\\:\\dots\\:,{X}_{n}\\right\\}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2. Data Cleaning and Preprocessing\u003c/h2\u003e \u003cp\u003eHandling Missing Data: Let \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{X}_{missing}\\)\u003c/span\u003e\u003c/span\u003e be the set of features with missing values. For each feature, missing values were filled by the median \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{\\sim}{X},\\)\u003c/span\u003e\u003c/span\u003e calculated as:\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$$\\:\\stackrel{\\sim}{{X}_{i}}=median\\:\\left({X}_{i}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhich ensures that the central tendency of the data is preserved while handling the missing entries.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eFeature Selection\u003c/b\u003e: Not all features are useful for prediction. A subset of features \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left\\{{X}_{1},\\:{X}_{2},\\:\\dots\\:\\dots\\:.,{X}_{k}\\right\\}\\)\u003c/span\u003e\u003c/span\u003e was chosen based on domain knowledge. Irrelevant features [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e55\u003c/span\u003e], such as the street address, were removed to simplify the model.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eData Type Conversion\u003c/b\u003e: The \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:date\\)\u003c/span\u003e\u003c/span\u003e feature was converted to a numerical format, enabling the analysis of time-related trends.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.3. Exploratory Data Analysis (EDA)\u003c/h2\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003ePrice Distribution\u003c/b\u003e: The distribution of the target variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:y\\:\\left(house\\:price\\right)\\)\u003c/span\u003e\u003c/span\u003e was examined using a probability density function (PDF) [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e56\u003c/span\u003e] and visualized using histograms. The distribution of prices was right-skewed, indicating the presence of high-priced houses that can be considered outliers.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe PDF is defined as:\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$$\\:f\\left(y\\right)=\\frac{1}{N}+{\\sum\\:}_{i=1}^{N}\\:\\left(\\delta\\:(y-{y}_{i}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{N}\\)\u003c/span\u003e\u003c/span\u003e is the number of samples, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\delta\\:(y-{y}_{i)}\\)\u003c/span\u003e\u003c/span\u003e is the Dirac delta function at \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:y=\\:{y}_{i}\\)\u003c/span\u003e\u003c/span\u003e .\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eScatter Plots\u003c/b\u003e: The relationship between house price ​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{y}\\)\u003c/span\u003e\u003c/span\u003e and each feature \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{X}}_{\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e was visualized using scatter plots. For example, the relationship between ​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{y}\\)\u003c/span\u003e\u003c/span\u003e (price) and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{X}}_{3}\\)\u003c/span\u003e\u003c/span\u003e (square footage of living area) can be described as:\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003cdiv id=\"Equ6\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$$\\:y=f\\:\\left({X}_{3}\\right)$$\u003c/div\u003e \u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eCorrelation Matrix\u003c/b\u003e: A correlation matrix was calculated to understand the strength of linear relationships [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e57\u003c/span\u003e] between the variables \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{X}}_{\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e. The correlation between two variables \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{X}}_{\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{X}}_{\\varvec{j}}\\)\u003c/span\u003e\u003c/span\u003e is defined by the Pearson correlation coefficient:\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003cdiv id=\"Equ7\" class=\"Equation\"\u003e \u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$$\\:{\\rho\\:}_{ij}=\\:\\frac{cov\\:({X}_{i},{X}_{j})}{\\sigma\\:{X}_{i}\\sigma\\:{X}_{j}}$$\u003c/div\u003e \u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{c}\\varvec{o}\\varvec{v}\\:({\\varvec{X}}_{\\varvec{i}},{\\varvec{X}}_{\\varvec{j}})\\)\u003c/span\u003e\u003c/span\u003e is the covariance of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{X}}_{\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{X}}_{\\varvec{j}}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{\\sigma\\:}{\\varvec{X}}_{\\varvec{i}}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{\\sigma\\:}{\\varvec{X}}_{\\varvec{j}}\\)\u003c/span\u003e\u003c/span\u003e are their standard deviations. The correlation matrix was visualized as heatmap, revealing that features like \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{s}\\varvec{q}\\varvec{f}\\varvec{t}\\_\\varvec{l}\\varvec{i}\\varvec{v}\\varvec{i}\\varvec{n}\\varvec{g}\\)\u003c/span\u003e\u003c/span\u003e had a high positive correlation with price.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e3.4. Outlier Detection\u003c/h2\u003e \u003cp\u003eOutliers are extreme values of the target variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varvec{y}\\)\u003c/span\u003e\u003c/span\u003e\u003cb\u003e​\u003c/b\u003e or any feature \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\varvec{X}}_{\\varvec{i}}\\:\\)\u003c/span\u003e\u003c/span\u003ethat deviate significantly from the majority of the data. Mathematically, an outlier can be defined as any point where:\u003cdiv id=\"Equ8\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e\n$$\\:{y}_{i}\u0026gt;\\:{Q}_{3}+1.5\\times\\:IQR\\:or\\:\\:\\:\\:\\:\\:\\:\\:{y}_{i}\\:\u0026lt;\\:{Q}_{1}-1.5\\:\\times\\:IQR$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e8\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{Q}_{1}\\:\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{Q}_{3}\\:\\)\u003c/span\u003e\u003c/span\u003eare the first and third quartiles, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:IQR=\\:{Q}_{3}-\\:{Q}_{1}\\:\\)\u003c/span\u003e\u003c/span\u003e is the interquartile range. Outliers were visually identified using box plots and scatter plots, particularly in relation to the price.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.5. Model Building and Evaluation\u003c/h2\u003e \u003cp\u003eA predictive model was developed to estimate house prices based on selected features. The target variable is denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:y,\\:\\)\u003c/span\u003e\u003c/span\u003eand the input feature vector is \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:X=\\left\\{{X}_{1},\\:{X}_{2},\\:\\dots\\:\\dots\\:.,{X}_{k}\\right\\}\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e3.5.1. Linear Regression Model\u003c/h2\u003e \u003cp\u003eA linear regression model was chosen for its simplicity. The model assumes a linear relationship between the target variable y and the input features ​\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{X}_{i}\\)\u003c/span\u003e\u003c/span\u003e, modeled as:\u003cdiv id=\"Equ9\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ9\" name=\"EquationSource\"\u003e\n$$\\:y=\\:{\\beta\\:}_{0}+\\:{\\beta\\:}_{1}{X}_{1}+\\:{\\beta\\:}_{2}{X}_{2}+\\dots\\:+{\\beta\\:}_{k}{X}_{k}+\\:\\in\\:$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e9\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\beta\\:}_{0}\\)\u003c/span\u003e\u003c/span\u003e is the intercept, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\{\\beta\\:}_{0},\\:{\\beta\\:}_{1},\\dots\\:.,{\\beta\\:}_{k}\\}\\:\\)\u003c/span\u003e\u003c/span\u003ewere estimated by minimizing the sum of squared residuals:\u003cdiv id=\"Equ10\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ10\" name=\"EquationSource\"\u003e\n$$\\:\\begin{array}{c}min\\\\\\:\\beta\\:\\end{array}\\:\\sum\\:_{i=1}^{N}{({y}_{i}-\\widehat{{y}_{i\\:}})}^{2}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e10\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\widehat{{y}_{i\\:}}\\)\u003c/span\u003e\u003c/span\u003eis the predicted price for the i-th house, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{y}_{i}\\)\u003c/span\u003e\u003c/span\u003e is the actual price.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e3.5.2. Model Training\u003c/h2\u003e \u003cp\u003eThe dataset was split into a training set (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{X}_{train},{y}_{train})\\)\u003c/span\u003e\u003c/span\u003e and a training set (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{X}_{test},\\:{y}_{test})\\)\u003c/span\u003e\u003c/span\u003e, with 80% of the data used for training and 20% for testing. The model was trained using the training set.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4. Model Evaluation","content":"\u003cp\u003eThe performance of the model was evaluated using Root Mean Square Error (RMSE) and R-squared (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{R}^{2}\\)\u003c/span\u003e\u003c/span\u003e) metrics:\u003cdiv id=\"Equ11\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ11\" name=\"EquationSource\"\u003e\n$$\\:RMSE=\\:\\sqrt{\\frac{1}{N}\\sum\\:_{i=1}^{N}{\\left({y}_{i}-\\widehat{{y}_{i\\:}}\\right)}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e11\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ12\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ12\" name=\"EquationSource\"\u003e\n$$\\:{R}^{2}=1-\\:\\frac{\\sum\\:_{i=1}^{N}{\\left({y}_{i}-\\widehat{{y}_{i\\:}}\\right)}^{2}}{\\sum\\:_{i=1}^{N}{\\left({y}_{i}-\\stackrel{-}{y}\\right)}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e12\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{y}\\:\\)\u003c/span\u003e\u003c/span\u003eis the mean of the observed prices. The RMSE provides an estimate of the average deviation of the predicted prices from the actual values, while \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{R}^{2}\\)\u003c/span\u003e\u003c/span\u003e represents the proportion of variance in \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:y\\)\u003c/span\u003e\u003c/span\u003e explained by the model.\u003c/p\u003e"},{"header":"5. Results and Discussions","content":"\u003cp\u003eThe analysis of the house price dataset revealed significant trends and relationships between various features and house prices. Through rigorous data preprocessing, exploratory data analysis (EDA), and the development of a linear regression model, key insights emerged that contribute to understanding the factors influencing housing prices in the market.\u003c/p\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e5.1. House Price Distribution\u003c/h2\u003e \u003cp\u003eThe initial exploratory analysis highlighted that house prices were predominantly right-skewed, with a concentration of properties priced between \u003cspan\u003e$\u003c/span\u003e200,000 and \u003cspan\u003e$\u003c/span\u003e500,000. The histogram depicting this distribution illustrated that while most transactions occurred in the mid-range, the presence of luxury homes significantly affected the overall average price. This skewness indicates that while affordable housing remains prevalent, high-end properties create a disparity in the perceived average market value.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e5.2. Correlation Analysis\u003c/h2\u003e \u003cp\u003eThe correlation matrix identified the square footage of living space and the quality grade of homes as the strongest predictors of house prices. With correlation coefficients of 0.70 and 0.66, respectively, these features demonstrated a clear relationship where larger and higher-quality homes were associated with increased prices. This finding aligns with existing literature, which suggests that buyers prioritize size and quality when evaluating property value.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e5.3. Correlation Heatmap\u003c/h2\u003e \u003cp\u003eThe heatmap displays the correlation coefficients ranging from \u0026minus;\u0026thinsp;1 to +\u0026thinsp;1, where:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eA value close to +\u0026thinsp;1 indicates a strong positive correlation,\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eA value close to -1 indicates a strong negative correlation,\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eA value around 0 suggests no correlation.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eIn this analysis, the heatmap highlights several significant correlations:\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eSquare Footage of Living Space (sqft_living): The highest positive correlation coefficient of 0.70 was observed, indicating that as the square footage increases, the price of the house tends to rise significantly. This emphasizes the importance of size in determining property values.\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eQuality Grade\u003c/strong\u003e \u003cp\u003eWith a correlation coefficient of 0.66, this feature also showed a strong positive relationship with house prices. Homes with higher quality grades are associated with higher market values, reflecting buyers' preferences for well-constructed and aesthetically appealing properties.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eNumber of Bathrooms\u003c/strong\u003e \u003cp\u003eThis feature exhibited a moderate positive correlation of 0.52, suggesting that more bathrooms generally contribute to higher prices, although the relationship is not as strong as that seen with square footage and quality grade.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e5.4. Comparison of Models\u003c/h2\u003e \u003cp\u003eThe comparison of models [\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e58\u003c/span\u003e] highlights the strengths of Linear Regression in house price prediction. Despite its simplicity, Linear Regression demonstrates competitive performance, offering a balance between accuracy and interpretability. Unlike complex models such as Gradient Boosting or Random Forest, Linear Regression provides clear insights into the relationship between features and target variables, making it more suitable for applications where transparency is crucial. Furthermore, Linear Regression is computationally efficient, requiring\u003c/p\u003e \u003cp\u003eless time and resources compared to tree-based models, which are prone to overfitting without careful tuning. While advanced models might slightly improve accuracy, the simplicity, speed, and ease of implementation of Linear Regression make it a reliable and practical choice for real-world applications, particularly when interpretability and efficiency are prioritized.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of Models.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026times;\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMAE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLinear Regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e210,908.173\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026times;\" colname=\"c3\"\u003e \u003cp\u003e9.869 \u0026times; 10\u003csup\u003e11\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.032284\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision Tree\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e262,910.016\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026times;\" colname=\"c3\"\u003e \u003cp\u003e1.052 \u0026times; 10\u003csup\u003e12\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e-0.031647\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e208,109.707\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026times;\" colname=\"c3\"\u003e \u003cp\u003e9.917 \u0026times; 10\u003csup\u003e11\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.027524\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boosting\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e202,521.382\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026times;\" colname=\"c3\"\u003e \u003cp\u003e9.814 \u0026times; 10\u003csup\u003e11\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.037695\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e5.5. Implications of Findings\u003c/h2\u003e \u003cp\u003eThe results of this study have important implications for both homebuyers and real estate professionals. Based on the regression coefficients, the following points can be highlighted:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eBedrooms: The negative coefficient for bedrooms (-66999.17) suggests that, all else being equal, an increase in the number of bedrooms tends to decrease the house price. This could indicate that buyers prioritize other factors, such as living space, over the number of bedrooms.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eLinear Regression Coefficients.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFeature\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCoefficient\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ebedrooms\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e-66999.169310\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ebathrooms\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e-7519.141800\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esqft_living\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e311.015539\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esqft_lot\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e-0.598012\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003efloors\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e30932.522099\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003econdition\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e61097.200192\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eBathrooms: The negative coefficient for \u003cem\u003ebathrooms\u003c/em\u003e (7519.14) implies that additional bathrooms may not significantly increase the house price, or that buyers may perceive diminishing returns on extra bathroom features.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSquare Footage of Living Space (\u003cem\u003esqft living\u003c/em\u003e): The positive coefficient for \u003cem\u003esqft living\u003c/em\u003e (311.02) indicates that larger living spaces contribute positively to house prices. This suggests that homebuyers value more living space, and homes with larger square footage tend to command higher prices.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eLot Size (\u003cem\u003esqft lot\u003c/em\u003e): The negative coefficient for \u003cem\u003esqft lot\u003c/em\u003e (0.60) implies that a larger lot size may not have a strong impact on house prices, or that buyers are more focused on the house\u0026rsquo;s features rather than the size of the lot.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFloors: The positive coefficient for \u003cem\u003efloors\u003c/em\u003e (30932.52) suggests that houses with more floors are generally valued higher, indicating that multi-story homes may appeal more to buyers, especially if they provide more living space.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCondition: The positive coefficient for \u003cem\u003econdition\u003c/em\u003e (61097.20) highlights that the condition of a home plays a significant role in determining its price. Homes in better condition tend to fetch higher prices, making it important for sellers to maintain the quality of their properties.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThese findings provide actionable insights for different stakeholders in the housing market:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eHomebuyers: Understanding these factors can help buyers prioritize their preferences. For instance, they may place greater importance on the size of the living space and the condition of the house, while less importance on the number of bedrooms and bathrooms.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eReal Estate Professionals: Agents can use these insights to adjust their marketing strategies, emphasizing properties with larger living spaces, better condition, and more floors as premium features.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePolicymakers: These results can guide policymakers in addressing housing affordability by focusing on mid-range housing options with better condition and reasonable sizes, which are likely to be in higher demand.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"6. Results and Discussions","content":"\u003cp\u003eThis analysis has highlighted the key factors influencing house prices, including the size of the living space and the quality of the home. Larger homes and those in better condition tend to be priced higher. The number of floors and the overall condition of the property also play significant roles in determining price, with better-maintained homes and those with more floors generally fetching higher values.\u003c/p\u003e \u003cp\u003eHowever, location was not fully addressed in this study, although it is clear that it plays a crucial role in pricing. The presence of high-value outliers in certain neighborhoods suggests that the area where a home is located can significantly impact its price. While the linear regression model worked well in capturing general trends, it may not be the best fit for predicting prices of luxury properties, which require more complex modeling.\u003c/p\u003e \u003cp\u003eIn the future, expanding the dataset to include economic factors and location-specific details, such as neighborhood amenities, could improve predictions. Advanced modeling techniques like decision trees or machine learning algorithms would also help capture more complex pricing patterns, particularly for luxury homes. This study has some limitations. The linear regression model, while effective, may not fully capture the complexities of luxury properties, which often have unique pricing patterns. Additionally, the dataset used is from one region, so future research could apply this model to different areas to see if the findings hold elsewhere.\u003c/p\u003e \u003cp\u003eThe study also didn\u0026rsquo;t consider factors like economic conditions or market changes over time. Including variables such as interest rates, inflation, or local economic data would provide a more complete understanding of house price dynamics. Future research should also explore more advanced techniques, like machine learning, to improve prediction accuracy and account for the unique features of high-end properties. Expanding the dataset and including more variables will help provide a deeper insight into the housing market and improve decision making for buyers, sellers, and policymakers.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgment\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe would like to thank our supervisor for his efforts and guidance.\u003c/p\u003e\n\u003ch3\u003eDeclaration of Competing Interests\u003c/h3\u003e\n\u003cp\u003eThe authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCRediT Authorship Contribution Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eXiaolin Ju:\u003c/strong\u003e Data curation, Software, Validation, Conceptualization, Methodology, Writing -review \u0026amp; editing, Investigation, Supervision. \u003cstrong\u003eVaskar Chakma:\u003c/strong\u003e Software, Conceptualization, Methodology, Writing -review \u0026amp; editing. \u003cstrong\u003eMisbahul Amin:\u003c/strong\u003e Software, Validation, Review \u0026amp; editing. \u003cstrong\u003eJoy Arkhid Chakma:\u003c/strong\u003e Validation, Review \u0026amp; editing.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSaunders C, Gammerman A, Vovk V (1998) Ridge regression learning algorithm in dual variables\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eForys I (2022) Machine learning in house price analysis: regression\u0026acute; models versus neural networks. Procedia Comput Sci 207:435\u0026ndash;445\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJordan MI, Mitchell TM (2015) Machine learning: Trends, perspectives, and prospects. Science 349(6245):255\u0026ndash;260\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePark B, Bae JK (2015) Using machine learning algorithms for housing price prediction: The case of fairfax county, virginia housing data. Expert Syst Appl 42(6):2928\u0026ndash;2934\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTruong Q, Nguyen M, Dang H, Mei B (2020) Housing price prediction via improved machine learning techniques. Procedia Comput Sci 174:433\u0026ndash;442\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNikou M, Mansourfar G, Bagherzadeh J (2019) Stock price prediction using deep learning algorithm and its comparison with machine learning algorithms. Intell Syst Acc Finance Manage 26(4):164\u0026ndash;174\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang Y-X, Hebert M (2016) Learning to learn: Model regression networks for easy small sample learning, in Computer Vision\u0026ndash;ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11\u0026ndash;14, 2016, Proceedings, Part VI 14, pp. 616\u0026ndash;634, Springer\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGong J, Sun S (2009) A new approach of stock price prediction based on logistic regression model, in 2009 International Conference on New Trends in Information and Service Science, pp. 1366\u0026ndash;1371, IEEE\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTaylor LO, Smith VK (2000) Environmental amenities as a source of market power. Land Econ, pp. 550\u0026ndash;568\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaulud D, Abdulazeez AM (2020) A review on linear regression comprehensive in machine learning. J Appl Sci Technol Trends 1(2):140\u0026ndash;147\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHimmelberg C, Mayer C, Sinai T (2005) Assessing high house prices: Bubbles, fundamentals and misperceptions. J Economic Perspect 19(4):67\u0026ndash;92\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi M (2015) Moving beyond the linear regression model: Advantages of the quantile regression model. J Manag 41(1):71\u0026ndash;98\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNimon KF, Oswald FL (2013) Understanding the results of multiple linear regression: Beyond standardized regression coefficients. Organizational Res Methods 16(4):650\u0026ndash;674\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHocking RR (2013) Methods and applications of linear models: regression and the analysis of variance. Wiley\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGhosh D, Vogt A, Outliers: An evaluation of methodologies, in Joint statistical meetings, vol. 12, pp. 3455\u0026ndash;3460, [16], Liu C, Xiong W (2012) China\u0026rsquo;s real estate market, 2018\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKahr J, Thomsett MC (2006) Real estate market valuation and analysis. Wiley\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHassan MM, Ahmad N, Hashim AH (2021) The conceptual framework of housing purchase decision-making process. Int J Acad Res Bus Social Sci 11(11):1673\u0026ndash;1690\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUllah F, Sepasgozar SM (2020) Key factors influencing purchase or rent decisions in smart real estate investments: A system dynamics approach using online forum thread data. Sustainability 12(11):4382\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuca JV, Muellbauer J, Murphy A (2021) What drives house price cycles? international experience and policy issues. J Econ Lit 59(3):773\u0026ndash;864\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChau KW, Chin T (2003) A critical review of literature on the hedonic price model. Int J Hous Sci Appl 27(2):145\u0026ndash;165\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen M-LT (2020) The hedonic pricing model applied to the housing market. Int J Econ Bus Adm 8(3):416\u0026ndash;428\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZaki J, Nayyar A, Dalal S, Ali ZH (2022) House price prediction using hedonic pricing model and machine learning techniques. Concurrency computation: Pract experience 34(27):e7342\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLisi G (2019) Property valuation: the hedonic pricing model\u0026ndash;location and housing submarkets. J Property Invest Finance 37(6):589\u0026ndash;596\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeyman AV, Law S, Berghauser Pont M (2018) How is location measured in housing valuation? a systematic review of accessibility specifications in hedonic price models. Urban Sci 3(1):3\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRosen S (1974) Hedonic prices and implicit markets: product differentiation in pure competition. J Polit Econ 82(1):34\u0026ndash;55\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRokach L, Maimon O (2005) Decision trees, Data mining and knowledge discovery handbook, pp. 165\u0026ndash;192\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Z (2021) Decision trees for objective house price prediction, in 2021 3rd International Conference on Machine Learning, Big Data and Business Intelligence (MLBDBI), pp. 280\u0026ndash;283, IEEE\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReddy PSM et al (2023) Decision tree regressor compared with random forest regressor for house price prediction in mumbai. J Surv Fisheries Sci 10(1):2323\u0026ndash;2332\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThamarai M, Malarvizhi S (2020) House price prediction modeling using machine learning. Int J Inform Eng Electron Bus, 12, 2\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRana VS, Mondal J, Sharma A, Kashyap I (2020) House price prediction using optimal regression techniques, in 2020 2nd International Conference on Advances in Computing, Communication Control and Networking (ICACCCN), pp. 203\u0026ndash;208, IEEE\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBreiman L (2001) Random forests. Mach Learn 45:5\u0026ndash;32\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHong J, Choi H, Kim W-s (2020) A house price valuation based on the random forest approach: the mass appraisal of residential property in south korea. Int J Strategic Property Manage 24(3):140\u0026ndash;152\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Y, Huang J, Zhang J, Liu S, Shorman S (2022) Analysis and prediction of second-hand house price based on random forest. 7(1):27\u0026ndash;42 Applied Mathematics and Nonlinear Sciences\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJamil SS, Bansal S, Vinjamuri L (2023) House price prediction using random forest techniques: a comparative study\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBentejac C, A. Cs\u0026acute; org\u0026uml; o, and, Mart˝ \u0026acute;ınez-Munoz G (2021) A comparative˜ analysis of gradient boosting algorithms, Artificial Intelligence Review, vol. 54, pp. 1937\u0026ndash;1967\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSibindi R, Mwangi RW, Waititu AG (2023) A boosting ensemble learning based hybrid light gradient boosting machine and extreme gradient boosting model for predicting house prices. Eng Rep 5(4):e12599\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHjort A, Pensar J, Scheel I, Sommervoll DE (2022) House price prediction with gradient boosted trees under different loss functions. J Property Res 39(4):338\u0026ndash;364\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi S, Jiang Y, Ke S, Nie K, Wu C (2021) Understanding the effects of influential factors on housing prices by combining extreme gradient boosting and a hedonic price model (xgboosthpm). Land 10(5):533\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang S, Li H, Li J, Zhang Y, Zou B (2018) Automatic analysis of lateral cephalograms based on multiresolution decision tree regression voting, Journal of healthcare engineering, vol. no. 1, p. 1797502, 2018\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeaman JA (2008) Black boxes. Emory LJ 58:427\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhan M, Debnath P, Al Sayeed A, Sumon MFI, Rahman A, Khan M, Pant L (2024) Explainable ai and machine learning model for california house price predictions: Intelligent model for homebuyers and policymakers. J Bus Manage Stud 6(5):73\u0026ndash;84\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeimer J, Francke M et al (2019) Out-of-sample house price prediction by hedonic price models and machine learning algorithms. Real Estate Res Q 18(2):13\u0026ndash;20\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMueller-Kett C (2024) Artificial intelligence for greater transparency in housing price estimation. AGILE: GIScience Ser 5:41\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eManjula R, Jain S, Srivastava S, Kher PR (2017) Real estate value prediction using multivariate regression models, in IOP Conference Series: Materials Science and Engineering, vol. 263, p. 042098, IOP Publishing\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMao Y, Yao R (2020) A geographic feature integrated multivariate linear regression method for house price prediction, in 2020 3rd international conference on humanities education and social sciences (ICHESS 2020), pp. 347\u0026ndash;351, Atlantis Press\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu S, Li Z, Qin Z, Yang X, Goh RSM (2017) A hybrid regression technique for house prices prediction, in IEEE international conference on industrial engineering and engineering management (IEEM), pp. 319\u0026ndash;323, IEEE, 2017\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu J (2022) Prediction on housing price based on the data on kaggle, in 2022 3rd International Conference on E-commerce and Internet Technology (ECIT 2022), pp. 627\u0026ndash;634, Atlantis Press\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen N (2022) House price prediction model of zhaoqing city based on correlation analysis and multiple linear regression analysis, Wireless Communications and Mobile Computing, vol. no. 1, p. 9590704, 2022\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYıldız S (2023) and P. Kara-\u0026uml; dayı Atas\u0026cedil;, A novel hybrid house price prediction model, Computational economics, vol. 62, no. 3, pp. 1215\u0026ndash;1232\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYahya N, Zainuddin NMM, Sjarif NNA, Azmi NFM (2020) Correlation analysis of factors affecting the prediction of price of terrace houses in penang, malaysia: A case study. Open Int J Inf 8(2):18\u0026ndash;39\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCabuk KS, Cengiz SK, Guler MG, Ozturk H, Efe AC, Ulas MG, Karademir FP (2023) Chasing the objective upper eyelid symmetry formula; r2, rmse, poc, mae and mse\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChicco D, Warrens MJ, Jurman G (2021) The coefficient of determination r-squared is more informative than smape, mae, mape, mse and rmse in regression analysis evaluation. Peerj Comput Sci 7:e623\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMalpezzi S (1999) A simple error correction model of house prices. J Hous Econ 8(1):27\u0026ndash;62\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBatory D (2005) Feature models, grammars, and propositional formulas, in International Conference on Software Product Lines, pp. 7\u0026ndash;20, Springer\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAfrasiabi M, Mohammadi M, Rastegar M, Stankovic L, Afrasiabi S, Khazaei M (2020) Deep-based conditional probability density function forecasting of residential loads. IEEE Trans Smart Grid 11(4):3646\u0026ndash;3657\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiles J, Shevlin M (2000) Applying regression and correlation: A guide for students and researchers\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSoegianto LM, Hinandra AT, Suri PA, Fajar M (2024) Comparison of model performance on housing business using linear regression, random forest regressor, svr, and neural network. Procedia Comput Sci 245:1139\u0026ndash;1145\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Nantong University","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"House Price Prediction, Linear Regression, Multivariate Analysis, Property Features, Market Valuation","lastPublishedDoi":"10.21203/rs.3.rs-5743165/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5743165/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis research examines the key factors that influence house prices, focusing on how size, condition, and structural features contribute to property valuation. A multivariate analysis using a Linear Regression model was conducted to assess the relationships between crucial features such as square footage, number of bedrooms, bathrooms, floors, and property condition. The analysis revealed that square footage and bathrooms exhibit the strongest positive correlations with house prices (both with correlation values of 0.76), indicating their significant impact on property valuation. In contrast, factors like condition and view demonstrated weaker correlations, suggesting a more limited influence. The Linear Regression model achieved an R-squared value of 0.75, explaining 75% of the variation in house prices based on these features. While the model effectively highlights key price determinants, its limitations in handling non-linear relationships and sensitivity to outliers are noted. This study emphasizes the importance of a nuanced, data-driven approach in understanding house price dynamics, offering valuable insights for buyers, sellers, and industry professionals. Future work could explore advanced predictive models and incorporate additional features to enhance forecasting accuracy.\u003c/p\u003e","manuscriptTitle":"What Drives House Prices? A Linear Regression Approach to Size, Condition, and Features","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-01-03 04:19:44","doi":"10.21203/rs.3.rs-5743165/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"f3b8f241-5ef8-487c-898e-4fd6b73b41be","owner":[],"postedDate":"January 3rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":42216610,"name":"Artificial Intelligence and Machine Learning"},{"id":42216611,"name":"Computer Architecture and Engineering"}],"tags":[],"updatedAt":"2025-01-03T04:19:44+00:00","versionOfRecord":[],"versionCreatedAt":"2025-01-03 04:19:44","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5743165","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5743165","identity":"rs-5743165","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-4.0