TFDF and TF-IDF in Financial Analysis

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

This paper develops a new text scoring heuristic for financial analysis, outperforming TF-IDF in trend prediction and precision, with further improvements using genetic algorithm optimization.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

The preprint studies how to weight and vectorize text from financial news for forecasting, comparing standard information-retrieval term weighting (especially TF-IDF) against a proposed “financial homegrown” heuristic. Using a news-text analysis framework focused on predicting trend direction and closeness of predicted values, the authors report that TF-IDF performs poorly for forecasting—faltering in trend prediction and yielding high errors on large numbers—while their new heuristic scores news items higher when they have higher publication/broadcast rates and yields better precision, with further gains after optimizing weights using a genetic-algorithm-based machine learning approach. A stated limitation is that the work is a preprint and not peer reviewed. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Textual analysis in the realm of business depends on text processing techniques borrowed mainly from information retrieval. However, it is not a viable solution capable of developing in finance. We suggest developing financial homegrown techniques for processing textual data. Especially in the course of scoring words where standard techniques are incongruous in financial analysis. On that matter, we pursue two issues. First , we examine major information retrieval heuristics. We find TF-IDF a facile solution that falters in predicting trend and generates high errors on large numbers. Second , we work on a new heuristic satisfying financial concernments. We consider the relation between publication rate of information and their importance. The proposed heuristic provides results of unmatchable performance in both predicting trend and precision measures. In additional analysis, we optimize our scheme using Genetic algorithm in a machine learning approach and get greater precision. In comparison with TF-IDF, the proposed heuristic conduces to 38.5 percent lower error in closeness measures which is again reduced by 16.46 percent with the help of machine learning. Our findings suggest that financial textual analysis and information retrieval should go their separate ways.
Full text 170,007 characters · extracted from preprint-html · click to expand
TFDF and TF-IDF in Financial Analysis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article TFDF and TF-IDF in Financial Analysis MEISAM HASHEMI, MEHRAN REZAEI, MARJAN KAEDI This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2883673/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Textual analysis in the realm of business depends on text processing techniques borrowed mainly from information retrieval. However, it is not a viable solution capable of developing in finance. We suggest developing financial homegrown techniques for processing textual data. Especially in the course of scoring words where standard techniques are incongruous in financial analysis. On that matter, we pursue two issues. First , we examine major information retrieval heuristics. We find TF-IDF a facile solution that falters in predicting trend and generates high errors on large numbers. Second , we work on a new heuristic satisfying financial concernments. We consider the relation between publication rate of information and their importance. The proposed heuristic provides results of unmatchable performance in both predicting trend and precision measures. In additional analysis, we optimize our scheme using Genetic algorithm in a machine learning approach and get greater precision. In comparison with TF-IDF, the proposed heuristic conduces to 38.5 percent lower error in closeness measures which is again reduced by 16.46 percent with the help of machine learning. Our findings suggest that financial textual analysis and information retrieval should go their separate ways. Financial textual analysis Term weighting Machine learning Stock market Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 1 INTRODUCTION Financial studies have benefited from analyzing news stories, tweets, comments, message boards, etc. using text processing techniques ( [ 1 ], [ 2 ], [ 3 ], [ 4 ]). However, they have overlooked developing or modifying text processing techniques compatible with the finance area. This gap motivates us to study text-processing technique in financial forecasting. Text-processing techniques are borrowed from information retrieval as well as from natural language processing areas. However, the basics of text analysis in financial forecasting and the two mentioned areas are not in concord. For example, the primary need in financial textual analysis is to discover the importance of documents whereas information retrieval primarily demands recovering relevant documents for an information need, and natural language processing pursuing the semantic and structural analysis of text. In each area, methods are developed to address their own problems that obviously cannot be equally successful in other areas. Therefore, financial textual analysis like other areas requires homegrown studies on text processing techniques. Every textual analysis needs a heuristic. The heuristic is supposed to discover certain information from statistical properties of words and documents. Given that, there is an incontrovertible relation between news events and market fluctuations [ 5 ], we suggest that the heuristic should be capable of discovering relational information between documents and prices. We have a novel approach to develop such a heuristic. We believe that important information has a high broadcast rate and impact on the market and hence can provide predictive information. Therefore, we work on a heuristic that can calculate higher scores for information with higher publication rate. For that matter, we diagnose and modify the standard weighting formula upon which we propose our heuristic. Financial forecasting considers two major issues. Closeness of predicted value and directional accuracy of predicted trend. We show that developing text-processing techniques meeting financial concernments is highly effective to reach greater precision in both value and trend prediction. In addition, our findings indicate that traditional techniques are inconsistent with financial forecasting and this area needs to work on such techniques as one of its own pursuits. The organization of this paper is such that the research problem and questions will be presented next. Related research and literature review are discussed in section 3. We explain our approach in section 4. Section 5 belongs to the evaluation design and in section 6, experimental results and discussion is presented. We conclude our work and suggest some thoughts for future works in section 7. 2 RESEARCH QUESTIONS Prior studies have demonstrated that (i) an ideal forecasting system should benefit from both data types, qualitative and quantitative, and (ii) qualitative data such as text can offer fair predictive information ( [ 6 ], [ 7 ], [ 1 ], [ 2 ]). Figure 1 sketches the general architecture of a financial forecasting system using text-mining techniques. The overall approach consists of both quantitative and qualitative data analysis components. The problem we seek a solution for is located in the qualitative data analysis part where text data are analyzed. Such data cannot be directly computed and then should be represented in computable form. The conversion of text is done in the section of weighting and feature vector creation. Our focus is to modify and improve this particular part, which is done traditionally using information retrieval techniques mostly TF-IDF. TF-IDF weighting scheme with value of \({W}_{i,j}\) for word \({W}_{i}\) in a news article \({d}_{j}\) is expressed in Eq. 1 , in which \({N}_{ }\) is the total number of news articles in the corpus, \({n}_{i}\) denotes the number of articles that contain word \(i\) and \({f}_{i,j}\) is the frequency of word \(i\) in document \(j\) [ 8 ]. $${W}_{i,j}= \left\{\begin{array}{c}\left(1+Log{f}_{i,j}\right)*Log\frac{N}{{n}_{i}} if {f}_{i,j}>0\\ 0 otherwise\end{array}\right.$$ 1 In textual analysis, the vogue among researchers is TF-IDF, though information retrieval has other heuristics on the roster. This scheme suggests that words appearing in fewer articles should get higher scores. Such heuristic gives documents containing least queried rare words a fair chance to be retrieved. Being rare for a word results in greater score and consequently more chance for the documents containing rare words to be ranked in a higher position. In information retrieval discipline, such a weighting mechanism makes sense. This leads to our first research question. How effective is the TF-IDF scheme in the financial context compared with other major information retrieval models? In news analysis, the weighting scheme is expected to reflect the importance of news stories. On the matter of importance, more important news is likely to have higher rate of broadcasting; results in a relationship suggesting the greater the importance, the higher the publication rate. Nevertheless, TF-IDF calculates lower scores for high broadcast news. Clearly, this quality of TF-IDF tends to be inherently problematic. “Applying nonbusiness word lists to accounting and finance topics can lead to a high misclassification rate and spurious correlations,” stated by Loughran and McDonald [ 9 ]. We follow up their conclusion and generalize it to nonbusiness text-processing techniques. We maintain that certain techniques should be developed especially for financial textual analysis. On that matter, we study a new term weighting technique capable of calculating higher scores for news with higher publication rate. This leads to our second research question. Do the news stories with higher publication rate contain more predictive information? Text processing leads to a high dimensional analysis that is a complex task. Even a perfectly contrived heuristic is not supposed to yield impeccable weights. We intend to optimize the computed weights and to fulfill the matter, we employ Genetic algorithm in a machine learning approach. 3 LITERATURE REVIEW Traders make decisions based on short-term and long-term strategies. The background philosophy comes from two different analyses, technical and fundamental. Technicians claim that a stock’s performance in the near future can be predicted by historical data. In such analysis, market volatility may be deemed as a promising condition. In contrast, fundamentalists make their investing decisions based on the intrinsic value of the companies. They consider basic economy factors and the companies’ financial statements. They balk at investing in a volatile market where share prices are not representing the value of the company. Share price even of a high valued company is in danger of tottering when fluctuations are escalating. A long-term investment made in a volatile market may face a huge fall and take quite a long period just to rebound. Predicting volatility in the market may suggest reconsidering investment decisions. For example, in a highly volatile period of Tehran stock market, from early 2014 to mid-2018, the price of even the most valued companies experiences considerable fall. Within the period, Mobarakeh Steel Company, which is regarded as one of the most reliable companies for long-term investment, takes a great fall from 5200 to 995 Rials. On the other hand, showing the stability of the market can support long-term investment decisions. Volatility is not always precluding investment. Technicians take extraordinary advantages of unstable market. Why so, as much as instability goes higher, traders may think of shorter periods of investment. For a highly volatile market where share prices experience a considerable amount of rise and fall in a day, technicians can make intraday trades. In addition, when the market sees weekly rises and falls, swinging strategy is the one to be considered. There are definite methods for technical and fundamental analysis but they are not capable of analyzing vast volumes of qualitative data such as news stories, tweets, comments and other useful textual data. In addition to such inability, Fama [ 6 ] and Malkiel [ 7 ] have introduced theories that nurture demand for analyzing textual data in financial forecasting. Fama [ 6 ] in Efficient Market Hypothesis clarifies that the price of a share is the reflection of all information related to that share. Next, Malkiel [ 7 ] in Random Walk Theory contends that it is impossible to predict share price effectively upon only historical data. They jointly allude to the usefulness of analyzing textual data, regardless of whether the analysis is technical or fundamental. We continue this section reviewing major information retrieval models. Next, belongs to a brief look at machine learning approaches and techniques used in financial textual analysis. We close this section by discussing data sources and word list that financial studies have benefited from or developed them. 3.1 Textual Representation The conversion of text data into computable form is done via textual representation process. The primary method for financial news articles is bag of words. This method comprises a set of processes such as text tokenization, stop-words removal and term weighting. In this method, each article is divided into words, as tokenization process. Then, a list of words with no meaning will be removed as stop-words removal phase and the remaining terms will be assigned scores as weighting process. The scores are obtained using a term weighting technique. Most weighting techniques come from information retrieval models such as Boolean, Probabilistic and Vector Space models [ 8 ]. The mostly used weighting technique is TF-IDF that is adapted from the Vector Space model and described by Eq. 1 . Equation 2 expresses the weighting technique used in Boolean model [ 8 ]. In this equation, the inquired word is shown by \(c\left(q\right)\) and \(c\left({d}_{j}\right)\) is the document j . This scheme gives one to the existing words in the document j and zero to those that do not exist. $${w}_{i,j}= \left\{\begin{array}{c}1 if \exists c\left(q\right) | c\left(q\right)=c({d}_{j})\\ 0 otherwise \end{array}\right.$$ 2 Probabilistic term weighting technique is characterized by Eq. 3 [ 8 ]. In this formula, \({P}_{i}\) indicates the probability of existing word i in the set of documents R . $${w}_{i,j}= \frac{{P}_{i}R}{1- {P}_{i}R}$$ 3 In a research on foreign exchange market using news analysis, Semiromi et al. [ 10 ] utilize the TF part of TF-IDF for weighting features. Their proposed model truly benefits from TF weighting and surprisingly delivers superior results – despite the fact that it is not a common weighting tactic. In a similar approach Hashemi et al. [ 11 ] use relative distribution of words instead of information retrieval heuristics in their stock market analysis using textual data and report greater results in precision measures. Findings of mentioned works concur to some extent with the basics of our approach. There are other methods seldom used in textual analysis such as noun phrases, named entities and n-gram analysis but they do not offer any discrete heuristic. Noun phrases and named entities are subsequent to bag of words that uses TF-IDF heuristic. They are supposed to improve the semantic and syntactic aspects of bag of words. Schumaker and Chen [ 2 ] propose a forecasting system using bag of words upon which they develop named entities, noun phrases and proper nouns as the subset of noun phrases. Butler and Kešelj [ 12 ] use n-gram analysis in terms of character-gram and word-gram to analyze companies reports in their trading engine. TF-IDF and other heuristics have been modified and improved in information retrieval studies but in textual analysis, the dominant version of TF-IDF is the original one. We implement and examine the prevailing form of techniques and compare them with our model and the enhanced approach, which uses machine learning. 3.2 Machine Learning Apart from algorithm, the common approach is classification. In this method, some classes such as positive and negative are defined. Then, historical data are used to identify trend and the classifier labels news articles in predefined classes. Support Vector Machine is the most common classification technique used by Antweiler and Frank [ 1 ], Schumaker and Chen [ 2 ] and Arias et al. [ 3 ]. The number of classes varies in different works; for example, Antweiler and Frank [ 1 ] and Mittermayer [ 13 ] use three, Seo et al. [ 14 ] use five and Thomas and Sycara [ 15 ] use two categories. In the matter of algorithm, there has been employed a variety of machine learning algorithms in financial studies. Thomas and Sycara [ 15 ] use Genetic Algorithm to analyze discussion board data using the number of messages and words posted per day and report of a maximum 30% return improvement. Arias et al. [ 3 ] use Decision Tree and Support Vector Machine in their predicting system, which they called summary tree , and state that in both areas of stock market and box office revenue their proposed approach can take advantage of textual analysis. Gidofalvi and Elkan [ 16 ] propose a text classifier based on Naïve Bayesian technique to classify news articles as up, down and unchanged according to the stock movement and find predictive possibility in a time frame of 20 minutes before and after a news article is published. Amin-Naseri and Gharacheh [ 17 ] use Genetic algorithm and artificial neural network in an attempt to predict long-term trends, which results in 78% average accuracy. Feuerriegel and Gordon [ 18 ] have employed decision support system to analyze news documents. They have achieved error reduction in RMSE of 19.5%, 19.4% and 35.6% for the DAX, CDAX and STOXX Europe 600 respectively. Furthermore, the other various techniques such as linear regression, logistic regression, k-means clustering have been used in fewer studies ( [ 19 ], [ 20 ], [ 21 ], [ 22 ] [ 23 ]). In practice, textual analysis turns into a multidimensional complex analysis that forces us to resort to machine learning. In prior financial textual studies, the main purpose of machine learning is to aid in directional forecasting. It also helps to reduce closeness error and to earn higher return in trading engines. However, we have a different approach. Our focus is to develop a financial related heuristic for weighting words but as it is not supposed to act impeccably, we intend to adjust words weight with the help of an optimization approach using Genetic algorithm. 3.3 Data Sources and Word Lists Textual data are generated by either users or companies, news outlets and wire services. User generated data such as comments, blog posts, tweets, message and discussion boards and bulletins are widely used in text mining. This type of data is very useful for some analysis such as opinion mining and sentiment analysis. News, quarterly and annual reports and even analysis by experts in form of articles can be recognized as official data. This type of data is devoid of users’ emotions and opinions and hence are not so beneficial in opinion mining and sentiment analysis. Message boards are used by Antweiler and Frank [ 1 ] in a classification system. They collect data from Yahoo Finance and Raging Bull message boards, in addition to news stories, and report of 1.5 million messages by the total number of words between 20 and 50 for the most part of the messages. They state that few messages have more than 200 words and the messages having more than 500 words are rare. Twitter data are one of the most used sources of textual data from which several studies have benefited such as Bollen et al. [ 4 ], Arias et al. [ 3 ], O’Connor et el. [ 24 ] and Culotta [ 21 ]. Bollen et al. [ 4 ] have studied tweets mood. They attempt to reveal users’ emotions from tweets as a kind of predictive information. They collect about 10 million tweets and develop a model named SOFNN to analyze the market. They report significant prediction improvement using their model and conclude that public calmness is predictive of the stock market. The very common source of data is news stories. Numerous studies have used news articles as data source ( [ 1 ], [ 2 ], [ 10 ], [ 13 ], [ 18 ], [ 23 ], [ 25 ], [ 26 ]). Schumaker and Chen [ 2 ] use news articles in their developed predicting machine named AZFinText system. They collect news data from Yahoo Finance for most of S&P 500 companies. They set a time constraint and just gather news released one hour after opening the market until 20 minutes before closing. Their dataset includes 9211 news articles, which are then filtered by some measures and finally results in 2839 articles for being used in bag of words. Mittermayer [ 13 ] works on a trading engine named NewsCATS and uses press releases. He sets a time restriction where press releases are excluded if they are published when the market is closed. The total number of press releases reaches 6602 in his work. He reports of a maximum 0.11% profit and outperforming the random trader. Harvard’s General Inquirer is a very common word list in the context of measuring expressive quality of text, emotionally categorized in more than 180 classes. However, it is not developed for the domain of finance and so has the potential to bring about some misclassification issues. Some researchers have decided to develop financial related word list in their studies. Loughran and McDonald [ 9 ] enunciate that about three-fourths of words tagged negative in Harvard Dictionary are not negative in the context of finance. They develop another negative word list and attempt to reveal the negative tone of textual data. In another study, Yu et al. [ 27 ] build two lexicons one for positive words and one for negative words. They use sentiment analysis to reveal emotional polarities toward special events from news stories. We examine at least 20 percent of all documents in our dataset and build a word list covering 263 important Persian words, considering Persian language limitations. In the review of previous works, one can conceive that many of them attempt to find answer for the question of whether any relation exists between market fluctuations and news stories, massage boards, tweets or comments or not. Many other studies focus on classifying articles for predicting trends or providing an adjustment to historical analysis. In this work, the major gap we study is term weighting that affects the whole system efficacy. We propose a new heuristic acceptable in finance and a Genetic-algorithm-based approach to enrich term weighting. To evaluate the proposed approach, we study the Tehran stock exchange using news stories from the central bank of Iran. 4 SCIENTIFIC APPROACH Figure 2 is the abstract design of the system we implement in this work. The word list alongside historical data and time-tagged news articles are sent to the model building unit. The word list is derived from news articles in the feature selection phase, which we explain in the following. To find answer for the first research question, in model building, we implement Probabilistic, Boolean and TF-IDF heuristics. Then, each entity of the word list is assigned a score, each time using a different heuristic. The only recourse to evaluate heuristics is to examine them in a forecasting system. So the weighted word list altogether news documents and historical data are then used in a forecasting system. We use the well-established evaluation criteria to examine error and performance of each model, which are explained in the next section. Textual analysis begins with preprocessing documents. One of the important sections in the preprocessing phase is feature selection. It relates to extracting the most beneficial words from a corpus. This section often is handled in an automatic way using statistical or semantic techniques. In addition, there are manual approaches done in terms of word sentiment. Although it may seem an ordinary routine, there are some intricate details. Schumaker and Chen [ 2 ] argue that the deficiency in bag of words model stems from a vast, unselective collection of words. It conveys that the preprocessing phase affects the ultimate outcome and also proves to a systematic handicap for the TF-IDF model. While evaluation should be done on equal terms, it can detract from our evaluation system. To overcome the issue, we handle this section manually and develop a word list. It ensures equality of assessment, as well as it makes the models performance independent from preprocessing issues. We carefully examine at least 20 percent of the whole news stories and glean influential words. The word list we develop covers 263 words that are considered important Persian terms in finance context. The recall score of our word list reaches 99.72 percent. 4.1 Weighting Features We believe that there is a causal relation between important information, traders’ behavior and market movements. Keynes [ 28 ] contends that experts are willing to analyze the behavior of the crowd of investors rather than financial numbers (as mentioned by Malkiel [ 7 ] p. 32). It implicitly alludes to the relation that the behavior of investors provides predictive information or actually forms market movements. Ozsoylev et al. [ 5 ] study investor networks where their findings reveal that every information event forms a mass behavior of investors; moreover, such information reported in news is linked with significant movement in stock market. We conclude that if important news galvanizes traders into action and market movements originate from traders’ actions then important news is the root of market movement and contains predictive information. This proposition has inspired us to work on a heuristic able to discover important information. As important information is supposed to be reported in several ways, the heuristic needs to seek information with a higher publication rate. A heuristic is a conjecture about the way data connect to each other. Every heuristic is designed to discover certain data that can be meaningful only for a specific purpose. Several heuristics with different tactics may be designed for a particular purpose but it is not reasonable to use a heuristic for a purpose different from what it has been designed for. Therefore, the heuristic should change if the purpose changes. Despite the fact that every purpose needs felicitous heuristic, TF-IDF is deemed as a panacea for all purposes. Some studies report that TF-IDF is effective in their analysis but the real truth is disclosed when studies compare TF-IDF with other methods. For example, Shumaker and Chen [ 2 ] report about 7 and 20 percent greater performance than TF-IDF for noun phrases and named entities respectively, in financial textual analysis. The hypothetical heuristic will give higher scores to information with higher publication rate that is interpreted as higher occurrence in the corpus. In contrast, TF-IDF is designed to give higher scores to the information with low occurrence in the dataset. The TF part, as term frequency, calculates score directly matching to word frequency in a particular document. The IDF part, as inverse document frequency, calculates scores directly opposite to the word occurrence among documents. The TF score is local, limited to a specific document, whereas the IDF score is general, and is the same in every document. Since important information contains words with high general occurrence among documents, the inversion of document frequency results in low weights for them. We suggest considering document frequency (DF) instead of inverse document frequency (IDF). In Eq. 4 , we invert IDF, on the left side, to DF, on the right side. By this modification, the higher weights will be associated with words appearing in more documents. We believe such heuristic is more congruent with what is truly happening in financial information dissemination and caters to our hypothesis. $${IDF}_{i}=Log\frac{N}{{n}_{i}} \underset{\Rightarrow }{Invert} {DF}_{i}=Log\frac{{n}_{i}}{N}$$ 4 The described change in IDF will reverse the TF-IDF weighting outcome. It converts TF-IDF into a new formula with an almost reversed heuristic, which we call TFDF. In practice, the number of appearance of words in documents is always less than the number of total documents in a dataset that makes the division of \(\frac{{n}_{i}}{\text{N}}\) to be less than one. Then, the DF part will bring negative weight for almost all words, probably except stop words. To avoid this issue, we add one to the result of division, before taking logarithm. Finally, by joining the modified IDF part, as DF, to TF part we can propose our weighting technique that is able to assign higher scores to more important news articles. Equation 5 presents the proposed method of weighting features in financial news context. In this equation \({n}_{i}\) is the number of documents containing word \(i\) and \(N\) denotes the total number of documents. $${W}_{i,j}= \left(1+Log{f}_{i,j}\right)\text{*}Log (1+ \frac{{n}_{i}}{\text{N}} )$$ 5 4.2 Weighting Enhancement In practical terms, textual analysis is a high dimensional data analysis in which each dimension represents a word. Term weighting scheme is a heuristic to calculate scores for dimensions. Finding the right score that can represent the true effect of dimension on the market is a complex problem. To overcome this problem, we follow up the term weighting adjustment with machine learning tactic. The problem we seek a solution for is about optimizing the outcome of term weighting technique. Therefore, we solve this problem using Genetic algorithm in a machine learning approach that is especially used for optimization problems. Learning section in Genetic algorithm tries to optimize chromosomes, or in depth genes of a chromosome. Every gene of a chromosome is in a relationship with the weight of a term that can be explained by summation, multiplication or other simple or complex formula. According to empirical findings, we have settled on summation where we get better outcomes from adding term weight with gene. Figure 3 displays the implemented Genetic algorithm. After textual representation, each news article results in a feature vector. They in addition to the historical data and selected terms will be used to compute chromosomes fitness. If any of the end-conditions is satisfied in a generation, the iteration of Genetic algorithm will be stopped and the most fitted member in the last generation will be selected as the optimum solution. Every chromosome is the subject of a fitness function and the result will be used as a tag of that chromosome. For as many as train set members (e.g., ‘n’) the process of fitness will run over every chromosome, hence there will be ‘n’ tags associated with a chromosome, and its final solution calculated as the mean of ‘n’ tags. Fitness value is calculated with absolute error. When fitness values are very close and the contrast between them needs to be magnified the choice will be squared error. The error is equal to distance of forecasted value and observed value. The average of these values will be taken to calculating general fitness via Eq. 6 . In this equation, \({A}_{i}\) and \({F}_{i}\) indicate observed and forecasted values in time of \(i\) respectively, and \(n\) is the total number of train set members. $$General Fitness = \frac{1}{n} \sum _{i=1}^{n}\left|{A}_{i}- {F}_{i}\right|$$ 6 There are several common techniques for the selection. In this work, we employ an elitist method with negligible chance and the intention is to increase the chance of selecting elites for the next generation. 5 EVALUATION DESIGN To evaluate weighting techniques in the context of finance the optimum way is to assess them in a prediction. Information retrieval has its own evaluation practices to determine the efficacy of a weighting scheme but they are widely off-topic methods in financial analysis. We assess the competence of all models by performing a long-term forecasting on Tehran stock exchange overall index. Tehran stock exchange has been active since 1967. During this period, 672 companies have been engaged and it has grown into a large market that recently attracted researchers’ attention ( [ 29 ], [ 30 ], [ 31 ], [ 32 ], [ 33 ], [ 34 ]). There are 13 main and 45 industry indices on Tehran stock exchange where the overall index is the most significant and referred one among investors. Overall index measures the performance of 367 major components. The performance of the overall index in reality is deemed by traders as the performance of the stock market and then they always keep the index movement under careful observation in their investments. A solid majority of companies that the overall index is under-effect their movements are governmental. Any qualitative data distributed from the government official media can affect their trends and hence the overall index. For example, news about any change in the exchange rate can bring about a huge effect on export-oriented companies; petroleum companies are under the effect of any change in the price of crack spreads or crude oil. All such and similar changes are issued and announced by the government itself. Many important governmental news is directly issued from the central bank media. The data set of this work includes financial news articles from the central bank of Iran and historical data from Tehran stock exchange from March 2005 until March of 2020. For evaluating the ultimate approach of this work, we utilized a function to cross validate test and train sets randomly and run it more than 10,000 times. From 178 months of collected data, 70% is chosen as train and the rest is used as test sets. Seeking answer for research questions, we examine qualitative and quantitative observations. First, we delve into predicted trends. We analyze information that can be helpful for different trading strategies offered by heuristics in trend prediction. Next, we discuss quantitative evaluation results obtained from mean absolute error (MAE), mean absolute percentage error (MAPE), root mean squared error (RMSE) and directional accuracy (DA), which are some of the well-established assessing techniques and expressed in equations 7 to 10 respectively. In these equations, \({E}_{t}\) and \({Y}_{t}\) denote forecasting error and observed value at time \(t\) , \(N\) is the total number of testing, \(TR and TF\) represent true directional forecasting and \(FR\) and \(FF\) represent false directional forecasting. $$MAE= \frac{1}{N} {\sum }_{t=1}^{N}\left|{E}_{t}\right|$$ 7 $$MAPE= \frac{100}{N} {\sum }_{t=1}^{N}\left|\frac{{E}_{t}}{{Y}_{t}}\right|$$ 8 $$RMSE= \sqrt{\frac{1}{n}{\sum }_{t=1}^{N}{E}_{t}^{2}}$$ 9 $$Directional Accuracy= \frac{TR+TF}{TR+TF+FR+FF}$$ 10 RMSE, as the root of MSE, and MSE both are useful to magnify the differences when errors are small or close together. On large errors, they generate very large numbers since they square the error. If a forecasting system gains small errors on average but rarely gains some large errors, then RMSE will be a very large number. Concretely, the general performance of a forecasting system cannot be fairly manifested by RMSE or MSE. In comparison, MAE and MAPE do not have bias on large numbers. Particularly MAPE that gauges the error relative to the actual value through dividing error by the observed value. Therefore, MAE and MAPE are preferable to determine the general performance . Moreover, Useful information can be deduced from RMSE together with MAPE. For example, low MAPE appeared with high RMSE can attest to a performance that generally gains low error while cannot equally well handle large numbers. In similar fashion, directional accuracy together with closeness measures provide deductive information. A forecasting method with large closeness error will predict the direction with large margin from the actual value. It becomes misleading when it comes to predicting direction in a monotonous market in which direction changes only with tiny amounts. In such a situation, predicting trend with a high margin from actual value gives spurious information that the market is volatile; as well as it may fail to detect major trend movements. In comparison, a prediction with small margin betokens a stable market – the information that those who are interested in buy and hold strategy are going after. 6 RESULTS AND DISCUSSION Given that we focus on evaluating the performance of heuristics, developing a competitive forecasting system is far from our scope. In our evaluation, forecasting is done merely using textual data, and results are pragmatic observations meant to assess and compare heuristics. They manifest the ability of heuristics in discovering predictive information. To answer the research questions, we delve into results and discuss experimental findings in terms of trend and error analysis. Magnified figures for discussions in section 6.1 are added in appendix. 6.1 Trend Analysis Figure 4 depicts observed and forecasted trends of Tehran stock exchange overall index. From 2005 until late 2009, the overall index is stable and consequently the forecasted values are relatively closer to the observed value. From late 2009 the trend gradually goes higher and then takes a sharp rise from mid-2012 to late 2013. From early 2014, it moves erratically until it starts a sharp rise in early 2018 again. Three considerable rising trends occur between the years of 2005 and 2020. First one starts in mid-2012 and moves the overall index from 23767 to 89723. The second and the smallest forms in late 2015 and lifts index from 61376 to 81723. This ends by the third stride that begins in early 2018, the index jumps from 92849 to 498700. In each specific period, there are possibilities in the form of risk and opportunity. We discuss supportive information that each model can offer. For the period 2005 to 2010, market moves almost on a horizontal line. Predicted trends by all models also attest to stability. In terms of predicting stability, the performance of all models are close. In late 2009 that the monotonous movement breaks, TF-IDF and TFDF methods are the first to notice that a nascent wave is emerging. When the market starts a more slope rise in mid-2012 it is revealed by TFDF, Probabilistic and Boolean, whereas TF-IDF detects a spurious oscillating movement. As a conclusion, in the period of 2005 to early 2014, all models perform closely in terms of predicting stability while in the matter of predicting incipient rising waves, which is useful for trend following strategy, TFDF is slightly better. There is fairly a lot error from early 2014 in all models however, predicted trends portend volatility that functions as a double-edged sword. Market instabilities are exaggerated by all models, highly by TF-IDF and slightly by TFDF. For those who are trading in short-terms with an eye toward technical analysis it signals a favorable opportunity that the market is swinging. They can benefit from recurrent rises and falls by making a swing-trading strategy. On the other hand, to those who are investing for long-term periods it indicates an inordinate amount of risk for buy and hold strategy. In such a situation, the market is not responding to fundamental data, which are the basis of buy and hold strategy. Therefore, they may better reconsider their investing decision or play it safe. The wavering movement breaks in late 2015 and starts a sharp rise where the TF-IDF model places above the both predicted and observed trends. Technically, TF-IDF is forecasting a false downward movement. The TFDF model can predict and reveal the rising trend clearly. Following TFDF, probabilistic and Boolean models detect the trend, and TF-IDF is last to do. Again, in early 2016 market movement reverts to its oscillating fashion where all models perform almost the same. For the period of early 2014 to mid-2018, we settle that all models are able to notify volatility albeit in different scales. In addition, the emerging ascend is clearly predicted by TFDF at the right time where other models do with a fair delay. After a lot of bounce for almost five years, since mid-2018, trend starts a big, steep rise. Remarkably, the TFDF model can predict the rise right before the starting time window of the sharp rise where TF-IDF fails and predicts a misleading fall. The Boolean and Probabilistic models are also able to reveal the start line of the trend though they fail to continue the trend. For the next rising trend, the first to detect movement is TF-IDF followed by TFDF. For the last trend, TFDF performs slightly better than other models. As a result, for the period of mid-2018 to early 2020, TFDF is able to predict all trends. Moreover, on the matter of time, TFDF reveals the major inchoate trend at the right time similar to Boolean and Probabilistic but more closely. TF-IDF shows weakness in predicting the rising trend at the propitious moment. As a conclusion, in predicting rising trends, the TFDF model can present a sound support for trend following decisions. In the matter of revealing volatility, we find that all models are able to do well but they predict fluctuations in different scales where traditional models exaggerate volatility more than our model. Most of the time, we see TF-IDF delivers a wavering trend, a signal which is translated to volatility, that is not a true barometer of actual trend. The information revealed by TFDF can boost technical based decisions in weekly and monthly investments. As well as it can warn fundamentalists about an unresponsive market to the intrinsic values, similar to the other models and on a more precise scale. 6.2 Error Analysis Figure 5 illustrates the absolute error of models in a logarithmic scale. Markers are representing the forecasting points. The amount of errors is depicted by how high the markers are above from the bottom. Therefore, the lower error leads to the lower place. We use logarithmic scale to avoid the high contrast map of errors that occurs in arithmetic scale. In the arithmetic scale, small errors and changes are shown in an imperceptible degree while large numbers are exceedingly magnified. The probabilistic model is placed above other models almost all the time and forms a line of circle markers on top. For most parts, it is like a competition for taking the lower places between Probabilistic and Boolean and for some parts, it occurs between TF-IDF and Probabilistic especially after mid-2018. It is even observable previously in Fig. 4 that the TF-IDF trend shows a higher buoyancy in that period. Surprisingly, the lowest points are registered by TF-IDF and Boolean models. From looking at the whole, the cloud of markers is demarcated by TFDF markers from the bottom, except some points. The line is not as solid as the one formed by the Probabilistic model. In some parts, there are competitions between most of the models to take lower places whereas for the most parts it is going on between TFDF and TF-IDF models. In conclusion, for most of time the figure shows that TF-IDF takes lower places than other traditional models, while it delineates a consistent tendency for TFDF to take lower places among all models. Table 1 displays the closeness of the predicted values. Smaller numbers in this table indicate more precision. The probabilistic model has the worst results in all measures and with a small amount Boolean model places in the next. While in RMSE, TF-IDF gains slightly higher error than Boolean, it reduces MAPE by 48 percent. RMSE is used to magnify the error to make sharp contrast in comparison. On the other side, MAPE is used to disclose the general error by calculating error relative to the actual value. In addition, prediction on large numbers is indeed conducive to large errors. Therefore, higher error in RMSE and lower error in MAPE can demonstrate that Boolean generates more error in general than TF-IDF. On the other hand, TF-IDF creates more error when it comes to predicting large values. Yet, it demonstrates that TF-IDF is much more reliable in general, relative to Probabilistic and Boolean. The lower absolute error in MAE by TF-IDF substantiates the suggestion. It may uphold the use of TF-IDF as the primary weighting method in textual analysis. Table 1 Mean of errors Method / Measures RMSE MAE MAPE Boolean 15418 11163 27.29 Probabilistic 18099 12939 31.06 TF-IDF 15739 9763 18.35 TFDF 10455 7109 15.59 Our model produces result of unmatchable precision in all measures. The lowest RMSE among traditional models is obtained by Boolean and TFDF lowers it by 47 percent. In terms of absolute error, the lowest MAE is obtained by TF-IDF then TFDF reduces it by 37 percent. Again, TF-IDF has the lowest MAPE relative to Boolean and Probabilistic but TFDF performs superiorly and achieves 17 percent more precision by reducing MAPE from 18.35 to 15.59. Directional accuracy is summarized in Table 2 . Theoretically, it is possible to guess the direction of movement 50 percent true only by chance. However, it may not be achievable randomly in the real world. Why so, the occurrence of rise and fall has never been evenly distributed. In practice, trends have bias and in the case of overall index, we observe the buoyancy of the Tehran stock exchange. Therefore, the thought of achieving 50 percent by chance is not a just consideration. Table 2 Directional accuracy Method / Measure Directional Accuracy Boolean 52.81 Probabilistic 51.12 TF-IDF 56.74 TFDF 57.30 By the first glance at the table, we see that the Probabilistic model acquires the least accuracy, as the same performance in closeness. TF-IDF performs the best among traditional models as the same act in absolute error. The Boolean model is not far ahead of probabilistic where it achieves 3 percent higher accuracy. TFDF presents greater results among all. Although it precedes TF-IDF only by a nominal margin, reliable performance in trend prediction reinforces its result. In contrast, poor performance in predicting trend blunts good result of TF-IDF. 6.3 Answer to Research Questions Figure 6 reviews precision and accuracy results of all models relative to TF-IDF. Bars represent errors so the lower is the better, except in directional accuracy (DA). Relative performances shown in the chart are useful to analyze how a model acts compared with others. Pursuing answer for the first research question, we scale results relative to TF-IDF that makes its performance clearer. For example, it shows great performance of TF-IDF in MAPE where Probabilistic and Boolean gain 69 and 48 percent more error respectively. To answer the first research question, we investigate and discuss the forecasting results. First , we set the actual moves of the overall index and predicted trends side by side, from 2005 to 2020. In terms of forecasting trend, TF-IDF is able to reveal volatility; in a few cases detects rising trend; sometimes gives spurious signals; and for the most part its trend is oscillating. Consequently, the inconstant performance renders the predicted trend invalid. Second , we examine closeness measures data of all models. On the matter of closeness, which lower is better, we assess models using three measures. RMSE data indicate that TF-IDF is far better than Probabilistic although with a slight margin it falls behind the Boolean. We observe vivid error reduction in MAE where TF-IDF wins against other information retrieval models, by generating the lowest absolute error. TF-IDF achieves an unbeatable score in MAPE where it was able to decrease the lowest error dramatically, which was from Boolean, by 48 percent. Third , we assess the accuracy of predicting direction for which we gather data from monthly forecasting. In that measure, TF-IDF gets itself in front but it is unreliable due to wavering trend prediction. In conclusion, except for the weakness in trend prediction, TF-IDF evidences better performance in general and is the best among major information retrieval models. Comparative results of TFDF and TF-IDF are summarized in Fig. 7 relative to TFDF. To make improvements clearer we scale results relative to the proposed model. It is even obvious to the casual observer that how unproductive is TF-IDF compared with TFDF, in closeness measures. The most improvement occurs in RMSE and the least does in directional forecasting. Given that, high RMSE can betoken of weakness in handling large numbers, it shows that the proposed model improves performance on large numbers by 50 percent. To answer the second research question, we field our model in a truly equal race among all models. First , we juxtapose the observed and forecasted trends of the overall index by all models. We consider information that is useful for short-term and long-term investments. On that matter, we analyze the information proffered by all models for different trading strategies. TFDF is greatly able to detect rising trends, which is useful for technicians especially those who are interested in trend-following strategy. It is the first to reveal all major rising trends in their nascent state. For those who are trading in short-terms TFDF is able to signal the opportune moments when volatility is escalating. As well as for fundamentalists it is able to show when the market is responsive to intrinsic values and when the market is tempestuous and too risky for buy and hold strategy. The proposed model performs better than any other did in trend forecasting. Second , in forecasting values we gauge how close the predicted values are to the real values. We examine the distances between actual and predicted values and observe that TFDF provides most of the shorter distances. On average, TFDF reduces absolute error by 37 percent. In other measures such as RMSE, our model gains the lowest error and increases closeness by 50 percent. Likewise, the proposed model decreases TF-IDF error in MAPE, which was the finest, by 17 percent. In closeness measures, TFDF produces unsurpassed results and reduces error by 34.6 percent on average. Third , we compare directional accuracy by examining the predicted directions in all 178 months and find that TFDF delivers the highest accuracy. There is no big difference between TFDF and TF-IDF in directional accuracy raw data but reliable performance in trend prediction reinforces such data. As a conclusion, the TFDF model benefits from our heuristic and evidently supports our hypothesis that news stories with higher publication rate provide more predictive information. 6.4 Enhanced Model and Closing Notes In Table 3 , performance of TFDF is matched against when it is aided by machine learning optimization, in terms of closeness measures. Machine learning enhances TFDF in all ways. In RMSE, where we see the lowest enhancement, machine learning lowers error by 14.32 percent. Then, it reduces the amount of absolute error by 15.70 percent. The most enhancement occurs in MAPE measure, which shows 19.37 percent more precision. The lowest error reduction obtains in RMSE and the greater in MAPE that forms a pattern similar to TF-IDF performance. Previously observed, TF-IDF is beaten by Boolean in RMSE with a small margin while acts excellently in MAPE and shrinks Boolean error by 48 percent. It is sensible to say that Genetic optimization suffers from the same difficulty as we mentioned for TF-IDF in forecasting large numbers, though it happens to a very smaller degree. Table 3 Mean of errors Method / Measures RMSE MAE MAPE TFDF 10455 7109 15.59 Enhanced TFDF 9145 6144 13.06 Table 4 compares directional accuracy of both TFDF and the enhanced model with machine learning. Here we encounter unexpected results when TFDF is enhanced via Genetic algorithm. Despite the palpable error reduction in closeness measures, in directional forecasting the TFDF model is ahead, although by a meagre amount. We expect more accuracy as we get more precision but it does not happen nonetheless. Table 4 Directional accuracy Method / Measure Directional Accuracy TFDF 57.30 Enhanced TFDF 57.14 We have an assumption. When variables of a prediction are close and trend moves implicitly, directional forecasting becomes more challenging than when the market is explicitly rising or falling. We cross validate test and train sets randomly and run the test more than 10,000 times. It may accidentally happen that test sets are filled, more than its share, with times from periods when market has been in the doldrums that eventually costs the trivial loss in directional accuracy. Precision and accuracy results of TFDF are summarized and compared with results when weights are enhanced using machine learning in Fig. 8 . To make obtained achievements by machine learning clearer we scale results relative to the enhanced model. Except for directional forecasting, machine learning succeeds in improving performance. The most change occurs in MAPE that can attest to general improvement in precision. Other studies train machine often to help predict trend movement. They develop classification systems mainly with the help of support vector machine. We train machine to optimize the weight of words. Our approach delivers perceptible effects on precision. On average, it helps to achieve 16.46 percent more precision in predicting value. Unexpectedly, it faces a trivial loss in directional accuracy compared with TFDF, while it is still better than other models. There is a reasonable prospect that directional accuracy improves with the help of a classifier, which may be considered as a minor future work. Schumaker and Chen argue that the weakness of bag of words stems from a vast collection of various words that works as a real handicap (bag of words is analogous to the TF-IDF model in our work). They suggest the better performance by other models comes from truncated representation of documents. However, it is not a cogent argument in our work regarding two reasons. First , we select the most influential words from the corpus that precludes bag of words from suffering a vast collection of words. Second , all models use the same word list that makes the measurement equal in all common ways. Therefore, there is no noisy word list nor a real handicap that gives reasons for the weakness in bag words; rather the weakness is inherent in its procedures where our findings reveal that the deficiency lies in the course of weighting terms. In brief, all heuristics unveil their ability in search of predictive information. The Probabilistic model produces poorest results in directional accuracy also in closeness measures. Boolean acts better than TF-IDF in predicting trend and gains lowest RMSE among information retrieval models. TF-IDF brings the finest results among major information retrieval models. In MAPE and MAE it gains the lowest error but falters in trend prediction. It answers our first research question and can justify why TF-IDF has been reckoned the standard model in textual analysis. However, none of the borrowed heuristics from information retrieval shows satisfying performance in all measures that brings up the proposition: the finance area should consider study on text-processing techniques as one of its own pursuits. On that matter, we propose a financial related heuristic. Our model outmaneuvers all models in all measures. The TFDF achieves highest precision and directional accuracy. It is the only model that detects all major rising trends in their nascent state. According to the general rules of forecasting, all models conduct themselves reasonably; however, the superiority of TFDF demonstrates that our heuristic is more consistent with financial forecasting bylaws. It answers our second research question and confirms that news with higher publication rate provides more predictive information. 7 CONCLUSION AND FUTURE WORKS Among major information retrieval models, TF-IDF is known as the de facto standard scheme in textual analysis. Its heuristic is perfectly compatible with problems in its origin topic. As the first issue considered in this work, we examine major information retrieval models in order to find whether TF-IDF is efficient in resolving business issues and if it is better than other standard models. We find that information retrieval models do not afford forecasting system the ability to detect all major movements. TF-IDF often delivers belated signs of rising trends and falls behind Boolean in RMSE by a small margin. In other closeness measures, it surpasses all standard models by providing 31.53 percent more precision on average. In directional accuracy, the Probabilistic model has the poorest result and though TF-IDF is ahead of Boolean by 6.93 percent, its result is vitiated by wavering trend prediction. Consequently, Boolean is preferable in trend prediction and forecasting on large numbers whereas TF-IDF acts superiorly in general according to MAPE and MAE measures. Text-processing techniques have not been improved or modified in compliance with norms in the context of business. They are merely borrowed from other areas, chiefly from information retrieval. We believe that nonbusiness text-processing techniques are incompatible with financial forecasting and potential to bring unproductive weights in the course of scoring words. It motivates us to work on a new weighting heuristic especially devised for financial analysis, as the main issue we consider in this work. We develop our new weighting heuristic upon the hypothesis: news stories with higher publication rate contain more predictive information. We diagnose the TF-IDF heuristic and find the problematic part, then we modify it to calculate a higher score for more important information. We maintain that the “inversion of document frequency” tends to vitiate the term weighting process. The acute shortcoming becomes manifest in comparing the results of TF-IDF with our model as TFDF. The proposed model shows reliable performance in discovering rising trends where TF-IDF fails. In closeness measures, it reduces errors of TF-IDF by 38.5 percent on average. TF-IDF provides competitive performance in directional accuracy but it is still beaten by TFDF nonetheless. Overall, the proposed heuristic brings better performance in all measures and evidently supports our hypothesis. In additional analysis, to overcome the high dimensionality of textual analysis we suggest optimizing scores using Genetic algorithm. Although a trivial loss surfaces in directional accuracy, it does not cast a shadow on the efficiency of the approach. We achieve 19.37 percent improvement in MAPE whereas 15.70 and 14.32 percent progress in MAE and RMSE respectively. On average, the enhanced method reduces the closeness errors by 16.46 percent. Our findings show that information retrieval solutions are not meant to succeed in the realm of business. Our model is not infallible but its performance still discloses how abortive information retrieval models are in discovering predictive information. They are designed for retrieval purposes that narrows their ideal efficacy within the retrieval’s orbit. Our evaluation focuses on revealing models efficiency in term weighting. Pursuing that, we use a word list in order to bypass the issues consequent upon feature selection and to ensure that models are not encumbered with any unforeseen disadvantage in preprocessing procedures. Developing word list manually is not a time-effective approach so in practical terms especially when data are large and mounting up regularly it may be recommended to select features automatically. As a future work, we suggest running feature selection using automatic settings in order to examine how models perform this section, thereupon the entire analysis as well. Declarations Ethical Approval Not applicable Competing interests Not applicable Authors' contributions All authors have contributed to the writing and reviewing of the manuscript text and the level of contribution is according to the name list order maintained in the paper. Funding This research and publication was not supported by any grant. Availability of data and materials All the materials including codes, experimental road map, data collected are available and ready to be provided as requested by the reviewers. References W. Antweiler and M. Z. Frank, "Is All That Talk Just Noise? The Information Content of Internet Stock Message Boards," The Journal of finance, vol. 59, no. 3, pp. 1259-1294, 2004. R. P. Schumaker and H. Chen, "Textual Analysis of Stock Market Prediction Using Breaking Financial News: The AZFinText System," ACM Transactions on Information Systems (TOIS), vol. 27, no. 2, pp. 1-19, 2009. M. Arias, A. Arratia and R. Xuriguera, "Forecasting with twitter data," ACM Transactions on Intelligent Systems and Technology (TIST), vol. 5, no. 1, pp. 1-24, 2014. J. Bollen, H. Mao and X. Zeng, "Twitter mood predicts the stock market," Journal of Computational Science, vol. 2, no. 1, pp. 1-8, 2011. H. N. Ozsoylev, JohanWalden, M. D. Yavuz and R. Bildik, "Investor Networks in the Stock Market," The Review of Financial Studies, vol. 27, no. 5, pp. 1323-1366, 2014. E. F. Fama, The behavior of stock-market prices, vol. 38, no. 1, pp. 34-105, 1965. B. G. Malkiel, A Random Walk Down Wall Street, New York: WW Norton & Company, 1999. R. Baeza-Yates and B. Ribeiro-Neto, Modern information retrieval, New York: ACM press, 1999. T. Loughran and B. McDonald, "When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks," The Journal of finance, vol. 66, no. 1, pp. 35-65, 2011. H. N. Semiromi, S. Lessmann and W. Peters, "News will tell: Forecasting foreign exchange rates based on news story events in the economy calendar," The North American Journal of Economics and Finance, vol. 52, p. 101181, 2020. M. Hahsemi, M. Rezaei and M. Kaedi, "Textual analysis of central bank news in forecasting long-term trend of Tehran stock exchange index," Journal of Information and Communication Technology, vol. 12, no. 44, pp. 119-132, 2020. M. Butler and V. Kešelj, "Financial forecasting using character n-gram analysis and readability scores of annual reports," in Canadian Conference on Artificial Intelligence, Berlin, Heidelberg, 2009. M.-A. Mittermayer, "Forecasting intraday stock price trends with text mining techniques," in Proceedings of the 37th Hawaii International Conference on Social Systems, Big Island, HI, USA, 2004. Y.-W. Seo, J. A. Giampapa and K. Sycara, "Text Classification for Intelligent Portfolio Management," Robotics Institute, Carnegie Mellon University , Pittsburgh, PA, 2002. J. Thomas and K. Sycara, "Integrating genetic algorithms and text learning for financial prediction," in Genetic and Evolutionary Computation Conference (GECCO), 2000. G. Gidofalvi and C. Elkan, "Using news articles to predict stock price movements," Department of Computer Science and Engineering, University of California, San Diego, 2001. M. R. Amin-Naseri and E. A. Gharacheh, "A hybrid artificial intelligence approach to monthly forecasting of crude oil price time series," in 10th International Conference on Engineering Applications of Neural Networks, 2007. S. Feuerriegel and J. Gordon, "Long-term stock index forecasting based on text mining of regulatory disclosures," Decision Support Systems, vol. 112, pp. 88-97, 2018. A. Sitaram and B. A. Huberman, "Predicting the future with social media," in IEEE/WIC/ACM international conference on web intelligence and intelligent agent technology, Toronto, ON, Canada , 2010. A. Tumasjan, T. Sprenger, P. Sandner and I. Welpe, "Predicting elections with twitter: What 140 characters reveal about political sentiment," in Fourth International Conference on Weblogs and Social Media, Washington, DC, USA, 2010. A. Culotta, "Detecting influenza outbreaks by analyzing Twitter messages," 2010. V. Lampos, T. D. Bie and N. Cristianini, "Flu detector-tracking epidemics on Twitter," in Joint European conference on machine learning and knowledge discovery in databases, Berlin, Heidelberg, 2010. K.-G. Aase, "Text Mining of News Articles for Stock," Institutt for datateknikk og informasjonsvitenskap , Trondheim, Norway, 2011. B. O'Connor, R. Balasubramanyan, B. Routledge and N. Smith, "From tweets to polls: Linking text sentiment to public opinion time series," in Proceedings of the International AAAI Conference on Weblogs and Social Media, Washington, DC, 2010. G. G.-R. Wu, T. C.-T. Hou and J.-L. Lin, "Can economic news predict Taiwan stock market returns?," Asia Pacific management review, vol. 24, no. 1, pp. 54-59, 2019. A. S. A. Rahman and S. Abdul-Rahman, "Mining textual terms for stock market prediction analysis using financial news," in International Conference on Soft Computing in Data Science, Singapore, 2017. W.-B. Yu, B.-R. Lea and B. Guruswamy, "A Theoretic Framework Integrating Text Mining and Energy Demand Forecasting," IJEBM, vol. 5, no. 3, pp. 211-224, 2007. J. M. Keynes, The General Theory of Employment, Interest, and Money, Harcourt, 1936. P. Falinouss, "Stock trend prediction using news articles: a text mining approach," Lulea University of Technology, 2007. M. Fasanghari and G. A. Montazer, "Design and implementation of fuzzy expert system for Tehran Stock Exchange portfolio recommendation," Expert Systems with Applications, vol. 37, no. 9, pp. 6138-6147, 2010. J. Zahedi and M. M. Rounaghi, "Application of artificial neural network models and principal component analysis method in predicting stock prices on Tehran Stock Exchange," Physica A: Statistical Mechanics and its Applications, vol. 438, pp. 178-187, 2015. P. Shahrestani and M. Rafei, "The impact of oil price shocks on Tehran Stock Exchange returns: Application of the Markov switching vector autoregressive models," Resources Policy, vol. 65, p. 101579, 2020. R. Ramezanian, A. Peymanfar and S. B. Ebrahimi, "An integrated framework of genetic network programming and multi-layer perceptron neural network for prediction of daily stock return: An application in Tehran stock exchange market," Applied soft computing, vol. 82, p. 105551, 2019. A. H. Ghahfarrokhi and M. Shamsfard, "Tehran stock exchange prediction using sentiment analysis of online textual opinions," Intelligent Systems in Accounting, Finance and Management, vol. 27, no. 1, pp. 22-37, 2020. Additional Declarations No competing interests reported. Supplementary Files supplements.docx APPENDIX.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2883673","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":197151118,"identity":"e58775b8-ef36-4c98-8a25-363bd7c17d3d","order_by":0,"name":"MEISAM HASHEMI","email":"","orcid":"","institution":"University of Isfahan","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"MEISAM","middleName":"","lastName":"HASHEMI","suffix":""},{"id":197151121,"identity":"2a43350a-20d0-4822-b3d9-d2990d53f9d1","order_by":1,"name":"MEHRAN REZAEI","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA20lEQVRIie3RvQrCMBDA8QtCXCJdTxB9AiESCB1EX6UiODmLg0hd2klfwVdJCdglODvq5uBQF+nk9yAusW6C+Q+5DPebDsDl+uGY95z4eFUhUg2/JcCLrd1qhpXdLp+0ayKd6z2JfPBiRfTYQqQqC46rAZNmPfBJhIAmgMRYCaXIqWZyM5T8TmADkIR2Us6D84WJ5eFJGgUIhSRSjCMT2zvhH4mmpeps0WdohhKCNbKW6YV2kkbkmJ86XS82IstG03o91fpoI1B6fSkGt5sCECt411nhVZfL5fqrrrslRghq6B3VAAAAAElFTkSuQmCC","orcid":"","institution":"University of Isfahan","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"MEHRAN","middleName":"","lastName":"REZAEI","suffix":""},{"id":197151123,"identity":"127b0d17-72fb-4a66-966a-29069281d940","order_by":2,"name":"MARJAN KAEDI","email":"","orcid":"","institution":"University of Isfahan","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"MARJAN","middleName":"","lastName":"KAEDI","suffix":""}],"badges":[],"createdAt":"2023-05-02 01:14:12","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2883673/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2883673/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":36661756,"identity":"2eaf5b1a-4d89-472c-a486-bee1c1724b50","added_by":"auto","created_at":"2023-05-05 22:17:55","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":121213,"visible":true,"origin":"","legend":"\u003cp\u003eGeneral architecture\u003c/p\u003e","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/b4c2a76f72d4f6471bd7371a.png"},{"id":36660322,"identity":"4ed03caa-9d5a-475e-8c74-1eea5aeebc63","added_by":"auto","created_at":"2023-05-05 21:53:55","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":75159,"visible":true,"origin":"","legend":"\u003cp\u003eSystem Architecture\u003c/p\u003e","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/543f2e55405c3dca12bb8ab4.png"},{"id":36660325,"identity":"5c3354f9-ae77-4dd7-a366-b9c44cdc251c","added_by":"auto","created_at":"2023-05-05 21:53:55","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":129811,"visible":true,"origin":"","legend":"\u003cp\u003eGenetic algorithm\u003c/p\u003e","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/89547f1843be577e50b723cb.png"},{"id":36660331,"identity":"f89d5623-6286-47a0-bda8-6dc9ddbd8a10","added_by":"auto","created_at":"2023-05-05 21:53:55","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":165727,"visible":true,"origin":"","legend":"\u003cp\u003eForecasting long-term\u003c/p\u003e","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/9aa3c84be4ccc53cdfbaaf30.png"},{"id":36660329,"identity":"789a6af1-a20d-48bc-bfcb-375725a646f6","added_by":"auto","created_at":"2023-05-05 21:53:55","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":282339,"visible":true,"origin":"","legend":"\u003cp\u003eThe Absolute errors in logarithmic scale\u003c/p\u003e","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/3d05d9da6ea7127f15d6a7e3.png"},{"id":36660327,"identity":"dd7f200b-013e-43a0-944a-a810973dc133","added_by":"auto","created_at":"2023-05-05 21:53:55","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":154325,"visible":true,"origin":"","legend":"\u003cp\u003ePrecision and Accuracy\u003c/p\u003e","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/b40042a0cb9d7776b401f880.png"},{"id":36661627,"identity":"2890cc05-4486-4335-b4b7-a2c488640b69","added_by":"auto","created_at":"2023-05-05 22:09:55","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":83951,"visible":true,"origin":"","legend":"\u003cp\u003ePrecision and Accuracy\u003c/p\u003e","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/11807a07c63bdd05f765d140.png"},{"id":36660962,"identity":"3b53cde9-3754-4799-b924-fe8d0647f61f","added_by":"auto","created_at":"2023-05-05 22:01:55","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":94385,"visible":true,"origin":"","legend":"\u003cp\u003ePrecision and Accuracy\u003c/p\u003e","description":"","filename":"Onlinefloatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/3423dc74c430f3b34e86d75f.png"},{"id":36733167,"identity":"081c5c81-6c9b-4746-bb49-7fb5a867732e","added_by":"auto","created_at":"2023-05-09 08:44:43","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1097562,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/ce16f771-224a-4950-be72-c79e044bfddc.pdf"},{"id":36660959,"identity":"ba118aa8-0875-4d29-8dfb-6d292367451c","added_by":"auto","created_at":"2023-05-05 22:01:55","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":13240,"visible":true,"origin":"","legend":"","description":"","filename":"supplements.docx","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/57679e5fd4a3d1a5826aa78d.docx"},{"id":36661625,"identity":"9b770eb3-3531-498a-b239-30c814a64ce1","added_by":"auto","created_at":"2023-05-05 22:09:55","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":31114,"visible":true,"origin":"","legend":"","description":"","filename":"APPENDIX.docx","url":"https://assets-eu.researchsquare.com/files/rs-2883673/v1/4c2b25c85888b51f03fadc50.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"TFDF and TF-IDF in Financial Analysis","fulltext":[{"header":"1 INTRODUCTION","content":"\u003cp\u003eFinancial studies have benefited from analyzing news stories, tweets, comments, message boards, etc. using text processing techniques ( [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e], [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]). However, they have overlooked developing or modifying text processing techniques compatible with the finance area. This gap motivates us to study text-processing technique in financial forecasting.\u003c/p\u003e \u003cp\u003eText-processing techniques are borrowed from information retrieval as well as from natural language processing areas. However, the basics of text analysis in financial forecasting and the two mentioned areas are not in concord. For example, the primary need in financial textual analysis is to discover the importance of documents whereas information retrieval primarily demands recovering relevant documents for an information need, and natural language processing pursuing the semantic and structural analysis of text. In each area, methods are developed to address their own problems that obviously cannot be equally successful in other areas. Therefore, financial textual analysis like other areas requires homegrown studies on text processing techniques.\u003c/p\u003e \u003cp\u003eEvery textual analysis needs a heuristic. The heuristic is supposed to discover certain information from statistical properties of words and documents. Given that, there is an incontrovertible relation between news events and market fluctuations [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], we suggest that the heuristic should be capable of discovering relational information between documents and prices.\u003c/p\u003e \u003cp\u003eWe have a novel approach to develop such a heuristic. We believe that important information has a high broadcast rate and impact on the market and hence can provide predictive information. Therefore, we work on a heuristic that can calculate higher scores for information with higher publication rate. For that matter, we diagnose and modify the standard weighting formula upon which we propose our heuristic.\u003c/p\u003e \u003cp\u003eFinancial forecasting considers two major issues. Closeness of predicted value and directional accuracy of predicted trend. We show that developing text-processing techniques meeting financial concernments is highly effective to reach greater precision in both value and trend prediction. In addition, our findings indicate that traditional techniques are inconsistent with financial forecasting and this area needs to work on such techniques as one of its own pursuits.\u003c/p\u003e \u003cp\u003eThe organization of this paper is such that the research problem and questions will be presented next. Related research and literature review are discussed in section 3. We explain our approach in section 4. Section 5 belongs to the evaluation design and in section 6, experimental results and discussion is presented. We conclude our work and suggest some thoughts for future works in section 7.\u003c/p\u003e"},{"header":"2 RESEARCH QUESTIONS","content":"\u003cp\u003ePrior studies have demonstrated that (i) an ideal forecasting system should benefit from both data types, qualitative and quantitative, and (ii) qualitative data such as text can offer fair predictive information ( [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e], [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]). Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e sketches the general architecture of a financial forecasting system using text-mining techniques. The overall approach consists of both quantitative and qualitative data analysis components. The problem we seek a solution for is located in the qualitative data analysis part where text data are analyzed. Such data cannot be directly computed and then should be represented in computable form. The conversion of text is done in the section of weighting and feature vector creation. Our focus is to modify and improve this particular part, which is done traditionally using information retrieval techniques mostly TF-IDF.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTF-IDF weighting scheme with value of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{i,j}\\)\u003c/span\u003e\u003c/span\u003e for word \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({W}_{i}\\)\u003c/span\u003e\u003c/span\u003e in a news article \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({d}_{j}\\)\u003c/span\u003e\u003c/span\u003e is expressed in Eq.\u0026nbsp;\u003cspan refid=\"Equ1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, in which \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({N}_{ }\\)\u003c/span\u003e\u003c/span\u003eis the total number of news articles in the corpus, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({n}_{i}\\)\u003c/span\u003e\u003c/span\u003e denotes the number of articles that contain word \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(i\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({f}_{i,j}\\)\u003c/span\u003e\u003c/span\u003e is the frequency of word \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(i\\)\u003c/span\u003e\u003c/span\u003e in document \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(j\\)\u003c/span\u003e\u003c/span\u003e [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$${W}_{i,j}= \\left\\{\\begin{array}{c}\\left(1+Log{f}_{i,j}\\right)*Log\\frac{N}{{n}_{i}} if {f}_{i,j}\u0026gt;0\\\\ 0 otherwise\\end{array}\\right.$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIn textual analysis, the vogue among researchers is TF-IDF, though information retrieval has other heuristics on the roster. This scheme suggests that words appearing in fewer articles should get higher scores. Such heuristic gives documents containing least queried rare words a fair chance to be retrieved. Being rare for a word results in greater score and consequently more chance for the documents containing rare words to be ranked in a higher position. In information retrieval discipline, such a weighting mechanism makes sense. This leads to our first research question.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eHow effective is the TF-IDF scheme in the financial context compared with other major information retrieval models?\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eIn news analysis, the weighting scheme is expected to reflect the importance of news stories. On the matter of importance, more important news is likely to have higher rate of broadcasting; results in a relationship suggesting the greater the importance, the higher the publication rate. Nevertheless, TF-IDF calculates lower scores for high broadcast news. Clearly, this quality of TF-IDF tends to be inherently problematic.\u003c/p\u003e \u003cp\u003e\u0026ldquo;Applying nonbusiness word lists to accounting and finance topics can lead to a high misclassification rate and spurious correlations,\u0026rdquo; stated by Loughran and McDonald [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. We follow up their conclusion and generalize it to nonbusiness text-processing techniques. We maintain that certain techniques should be developed especially for financial textual analysis. On that matter, we study a new term weighting technique capable of calculating higher scores for news with higher publication rate. This leads to our second research question.\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eDo the news stories with higher publication rate contain more predictive information?\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eText processing leads to a high dimensional analysis that is a complex task. Even a perfectly contrived heuristic is not supposed to yield impeccable weights. We intend to optimize the computed weights and to fulfill the matter, we employ Genetic algorithm in a machine learning approach.\u003c/p\u003e"},{"header":"3 LITERATURE REVIEW","content":"\u003cp\u003eTraders make decisions based on short-term and long-term strategies. The background philosophy comes from two different analyses, technical and fundamental. Technicians claim that a stock\u0026rsquo;s performance in the near future can be predicted by historical data. In such analysis, market volatility may be deemed as a promising condition. In contrast, fundamentalists make their investing decisions based on the intrinsic value of the companies. They consider basic economy factors and the companies\u0026rsquo; financial statements. They balk at investing in a volatile market where share prices are not representing the value of the company.\u003c/p\u003e \u003cp\u003eShare price even of a high valued company is in danger of tottering when fluctuations are escalating. A long-term investment made in a volatile market may face a huge fall and take quite a long period just to rebound. Predicting volatility in the market may suggest reconsidering investment decisions. For example, in a highly volatile period of Tehran stock market, from early 2014 to mid-2018, the price of even the most valued companies experiences considerable fall. Within the period, Mobarakeh Steel Company, which is regarded as one of the most reliable companies for long-term investment, takes a great fall from 5200 to 995 Rials. On the other hand, showing the stability of the market can support long-term investment decisions.\u003c/p\u003e \u003cp\u003eVolatility is not always precluding investment. Technicians take extraordinary advantages of unstable market. Why so, as much as instability goes higher, traders may think of shorter periods of investment. For a highly volatile market where share prices experience a considerable amount of rise and fall in a day, technicians can make intraday trades. In addition, when the market sees weekly rises and falls, swinging strategy is the one to be considered.\u003c/p\u003e \u003cp\u003eThere are definite methods for technical and fundamental analysis but they are not capable of analyzing vast volumes of qualitative data such as news stories, tweets, comments and other useful textual data. In addition to such inability, Fama [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] and Malkiel [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] have introduced theories that nurture demand for analyzing textual data in financial forecasting. Fama [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] in Efficient Market Hypothesis clarifies that the price of a share is the reflection of all information related to that share. Next, Malkiel [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] in Random Walk Theory contends that it is impossible to predict share price effectively upon only historical data. They jointly allude to the usefulness of analyzing textual data, regardless of whether the analysis is technical or fundamental.\u003c/p\u003e \u003cp\u003eWe continue this section reviewing major information retrieval models. Next, belongs to a brief look at machine learning approaches and techniques used in financial textual analysis. We close this section by discussing data sources and word list that financial studies have benefited from or developed them.\u003c/p\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Textual Representation\u003c/h2\u003e \u003cp\u003eThe conversion of text data into computable form is done via textual representation process. The primary method for financial news articles is bag of words. This method comprises a set of processes such as text tokenization, stop-words removal and term weighting. In this method, each article is divided into words, as tokenization process. Then, a list of words with no meaning will be removed as stop-words removal phase and the remaining terms will be assigned scores as weighting process. The scores are obtained using a term weighting technique. Most weighting techniques come from information retrieval models such as Boolean, Probabilistic and Vector Space models [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. The mostly used weighting technique is TF-IDF that is adapted from the Vector Space model and described by Eq.\u0026nbsp;\u003cspan refid=\"Equ1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eEquation \u003cspan refid=\"Equ2\" class=\"InternalRef\"\u003e2\u003c/span\u003e expresses the weighting technique used in Boolean model [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. In this equation, the inquired word is shown by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\left(q\\right)\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\left({d}_{j}\\right)\\)\u003c/span\u003e\u003c/span\u003e is the document \u003cem\u003ej\u003c/em\u003e. This scheme gives one to the existing words in the document \u003cem\u003ej\u003c/em\u003e and zero to those that do not exist.\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$${w}_{i,j}= \\left\\{\\begin{array}{c}1 if \\exists c\\left(q\\right) | c\\left(q\\right)=c({d}_{j})\\\\ 0 otherwise \\end{array}\\right.$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eProbabilistic term weighting technique is characterized by Eq.\u0026nbsp;\u003cspan refid=\"Equ3\" class=\"InternalRef\"\u003e3\u003c/span\u003e [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. In this formula, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({P}_{i}\\)\u003c/span\u003e\u003c/span\u003e indicates the probability of existing word \u003cem\u003ei\u003c/em\u003e in the set of documents \u003cem\u003eR\u003c/em\u003e.\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$${w}_{i,j}= \\frac{{P}_{i}R}{1- {P}_{i}R}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIn a research on foreign exchange market using news analysis, Semiromi et al. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] utilize the TF part of TF-IDF for weighting features. Their proposed model truly benefits from TF weighting and surprisingly delivers superior results \u0026ndash; despite the fact that it is not a common weighting tactic. In a similar approach Hashemi et al. [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] use relative distribution of words instead of information retrieval heuristics in their stock market analysis using textual data and report greater results in precision measures. Findings of mentioned works concur to some extent with the basics of our approach.\u003c/p\u003e \u003cp\u003eThere are other methods seldom used in textual analysis such as noun phrases, named entities and n-gram analysis but they do not offer any discrete heuristic. Noun phrases and named entities are subsequent to bag of words that uses TF-IDF heuristic. They are supposed to improve the semantic and syntactic aspects of bag of words. Schumaker and Chen [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] propose a forecasting system using bag of words upon which they develop named entities, noun phrases and proper nouns as the subset of noun phrases. Butler and Kešelj [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] use n-gram analysis in terms of character-gram and word-gram to analyze companies reports in their trading engine.\u003c/p\u003e \u003cp\u003eTF-IDF and other heuristics have been modified and improved in information retrieval studies but in textual analysis, the dominant version of TF-IDF is the original one. We implement and examine the prevailing form of techniques and compare them with our model and the enhanced approach, which uses machine learning.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Machine Learning\u003c/h2\u003e \u003cp\u003eApart from algorithm, the common approach is classification. In this method, some classes such as positive and negative are defined. Then, historical data are used to identify trend and the classifier labels news articles in predefined classes. Support Vector Machine is the most common classification technique used by Antweiler and Frank [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e], Schumaker and Chen [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] and Arias et al. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. The number of classes varies in different works; for example, Antweiler and Frank [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] and Mittermayer [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] use three, Seo et al. [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] use five and Thomas and Sycara [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] use two categories.\u003c/p\u003e \u003cp\u003eIn the matter of algorithm, there has been employed a variety of machine learning algorithms in financial studies. Thomas and Sycara [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] use Genetic Algorithm to analyze discussion board data using the number of messages and words posted per day and report of a maximum 30% return improvement. Arias et al. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] use Decision Tree and Support Vector Machine in their predicting system, which they called \u003cem\u003esummary tree\u003c/em\u003e, and state that in both areas of stock market and box office revenue their proposed approach can take advantage of textual analysis. Gidofalvi and Elkan [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] propose a text classifier based on Na\u0026iuml;ve Bayesian technique to classify news articles as up, down and unchanged according to the stock movement and find predictive possibility in a time frame of 20 minutes before and after a news article is published. Amin-Naseri and Gharacheh [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] use Genetic algorithm and artificial neural network in an attempt to predict long-term trends, which results in 78% average accuracy. Feuerriegel and Gordon [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] have employed decision support system to analyze news documents. They have achieved error reduction in RMSE of 19.5%, 19.4% and 35.6% for the DAX, CDAX and STOXX Europe 600 respectively.\u003c/p\u003e \u003cp\u003eFurthermore, the other various techniques such as linear regression, logistic regression, k-means clustering have been used in fewer studies ( [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]).\u003c/p\u003e \u003cp\u003eIn practice, textual analysis turns into a multidimensional complex analysis that forces us to resort to machine learning. In prior financial textual studies, the main purpose of machine learning is to aid in directional forecasting. It also helps to reduce closeness error and to earn higher return in trading engines. However, we have a different approach. Our focus is to develop a financial related heuristic for weighting words but as it is not supposed to act impeccably, we intend to adjust words weight with the help of an optimization approach using Genetic algorithm.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Data Sources and Word Lists\u003c/h2\u003e \u003cp\u003eTextual data are generated by either users or companies, news outlets and wire services. User generated data such as comments, blog posts, tweets, message and discussion boards and bulletins are widely used in text mining. This type of data is very useful for some analysis such as opinion mining and sentiment analysis. News, quarterly and annual reports and even analysis by experts in form of articles can be recognized as official data. This type of data is devoid of users\u0026rsquo; emotions and opinions and hence are not so beneficial in opinion mining and sentiment analysis.\u003c/p\u003e \u003cp\u003eMessage boards are used by Antweiler and Frank [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] in a classification system. They collect data from Yahoo Finance and Raging Bull message boards, in addition to news stories, and report of 1.5\u0026nbsp;million messages by the total number of words between 20 and 50 for the most part of the messages. They state that few messages have more than 200 words and the messages having more than 500 words are rare.\u003c/p\u003e \u003cp\u003eTwitter data are one of the most used sources of textual data from which several studies have benefited such as Bollen et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], Arias et al. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], O\u0026rsquo;Connor et el. [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] and Culotta [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Bollen et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] have studied tweets mood. They attempt to reveal users\u0026rsquo; emotions from tweets as a kind of predictive information. They collect about 10\u0026nbsp;million tweets and develop a model named \u003cem\u003eSOFNN\u003c/em\u003e to analyze the market. They report significant prediction improvement using their model and conclude that public calmness is predictive of the stock market.\u003c/p\u003e \u003cp\u003eThe very common source of data is news stories. Numerous studies have used news articles as data source ( [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e], [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e], [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e], [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]). Schumaker and Chen [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] use news articles in their developed predicting machine named \u003cem\u003eAZFinText\u003c/em\u003e system. They collect news data from Yahoo Finance for most of S\u0026amp;P 500 companies. They set a time constraint and just gather news released one hour after opening the market until 20 minutes before closing. Their dataset includes 9211 news articles, which are then filtered by some measures and finally results in 2839 articles for being used in bag of words. Mittermayer [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] works on a trading engine named \u003cem\u003eNewsCATS\u003c/em\u003e and uses press releases. He sets a time restriction where press releases are excluded if they are published when the market is closed. The total number of press releases reaches 6602 in his work. He reports of a maximum 0.11% profit and outperforming the random trader.\u003c/p\u003e \u003cp\u003eHarvard\u0026rsquo;s General Inquirer is a very common word list in the context of measuring expressive quality of text, emotionally categorized in more than 180 classes. However, it is not developed for the domain of finance and so has the potential to bring about some misclassification issues. Some researchers have decided to develop financial related word list in their studies. Loughran and McDonald [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] enunciate that about three-fourths of words tagged negative in Harvard Dictionary are not negative in the context of finance. They develop another negative word list and attempt to reveal the negative tone of textual data. In another study, Yu et al. [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e] build two lexicons one for positive words and one for negative words. They use sentiment analysis to reveal emotional polarities toward special events from news stories. We examine at least 20 percent of all documents in our dataset and build a word list covering 263 important Persian words, considering Persian language limitations.\u003c/p\u003e \u003cp\u003eIn the review of previous works, one can conceive that many of them attempt to find answer for the question of whether any relation exists between market fluctuations and news stories, massage boards, tweets or comments or not. Many other studies focus on classifying articles for predicting trends or providing an adjustment to historical analysis. In this work, the major gap we study is term weighting that affects the whole system efficacy. We propose a new heuristic acceptable in finance and a Genetic-algorithm-based approach to enrich term weighting. To evaluate the proposed approach, we study the Tehran stock exchange using news stories from the central bank of Iran.\u003c/p\u003e \u003c/div\u003e"},{"header":"4 SCIENTIFIC APPROACH","content":"\u003cp\u003eFigure \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e is the abstract design of the system we implement in this work. The word list alongside historical data and time-tagged news articles are sent to the model building unit. The word list is derived from news articles in the feature selection phase, which we explain in the following. To find answer for the first research question, in model building, we implement Probabilistic, Boolean and TF-IDF heuristics. Then, each entity of the word list is assigned a score, each time using a different heuristic. The only recourse to evaluate heuristics is to examine them in a forecasting system. So the weighted word list altogether news documents and historical data are then used in a forecasting system. We use the well-established evaluation criteria to examine error and performance of each model, which are explained in the next section.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTextual analysis begins with preprocessing documents. One of the important sections in the preprocessing phase is feature selection. It relates to extracting the most beneficial words from a corpus. This section often is handled in an automatic way using statistical or semantic techniques. In addition, there are manual approaches done in terms of word sentiment. Although it may seem an ordinary routine, there are some intricate details.\u003c/p\u003e \u003cp\u003eSchumaker and Chen [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] argue that the deficiency in bag of words model stems from a vast, unselective collection of words. It conveys that the preprocessing phase affects the ultimate outcome and also proves to a systematic handicap for the TF-IDF model. While evaluation should be done on equal terms, it can detract from our evaluation system. To overcome the issue, we handle this section manually and develop a word list. It ensures equality of assessment, as well as it makes the models performance independent from preprocessing issues.\u003c/p\u003e \u003cp\u003eWe carefully examine at least 20 percent of the whole news stories and glean influential words. The word list we develop covers 263 words that are considered important Persian terms in finance context. The recall score of our word list reaches 99.72 percent.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Weighting Features\u003c/h2\u003e \u003cp\u003eWe believe that there is a causal relation between important information, traders\u0026rsquo; behavior and market movements. Keynes [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] contends that experts are willing to analyze the behavior of the crowd of investors rather than financial numbers (as mentioned by Malkiel [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] p. 32). It implicitly alludes to the relation that the behavior of investors provides predictive information or actually forms market movements. Ozsoylev et al. [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] study investor networks where their findings reveal that every information event forms a mass behavior of investors; moreover, such information reported in news is linked with significant movement in stock market. We conclude that if important news galvanizes traders into action and market movements originate from traders\u0026rsquo; actions then important news is the root of market movement and contains predictive information.\u003c/p\u003e \u003cp\u003eThis proposition has inspired us to work on a heuristic able to discover important information. As important information is supposed to be reported in several ways, the heuristic needs to seek information with a higher publication rate.\u003c/p\u003e \u003cp\u003eA heuristic is a conjecture about the way data connect to each other. Every heuristic is designed to discover certain data that can be meaningful only for a specific purpose. Several heuristics with different tactics may be designed for a particular purpose but it is not reasonable to use a heuristic for a purpose different from what it has been designed for. Therefore, the heuristic should change if the purpose changes.\u003c/p\u003e \u003cp\u003eDespite the fact that every purpose needs felicitous heuristic, TF-IDF is deemed as a panacea for all purposes. Some studies report that TF-IDF is effective in their analysis but the real truth is disclosed when studies compare TF-IDF with other methods. For example, Shumaker and Chen [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] report about 7 and 20 percent greater performance than TF-IDF for noun phrases and named entities respectively, in financial textual analysis.\u003c/p\u003e \u003cp\u003eThe hypothetical heuristic will give higher scores to information with higher publication rate that is interpreted as higher occurrence in the corpus. In contrast, TF-IDF is designed to give higher scores to the information with low occurrence in the dataset. The TF part, as term frequency, calculates score directly matching to word frequency in a particular document. The IDF part, as inverse document frequency, calculates scores directly opposite to the word occurrence among documents. The TF score is local, limited to a specific document, whereas the IDF score is general, and is the same in every document. Since important information contains words with high \u003cem\u003egeneral occurrence\u003c/em\u003e among documents, the inversion of document frequency results in low weights for them.\u003c/p\u003e \u003cp\u003eWe suggest considering document frequency (DF) instead of inverse document frequency (IDF). In Eq.\u0026nbsp;\u003cspan refid=\"Equ4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, we invert IDF, on the left side, to DF, on the right side. By this modification, the higher weights will be associated with words appearing in more documents. We believe such heuristic is more congruent with what is truly happening in financial information dissemination and caters to our hypothesis.\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$${IDF}_{i}=Log\\frac{N}{{n}_{i}} \\underset{\\Rightarrow }{Invert} {DF}_{i}=Log\\frac{{n}_{i}}{N}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe described change in IDF will reverse the TF-IDF weighting outcome. It converts TF-IDF into a new formula with an almost reversed heuristic, which we call TFDF.\u003c/p\u003e \u003cp\u003eIn practice, the number of appearance of words in documents is always less than the number of total documents in a dataset that makes the division of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\frac{{n}_{i}}{\\text{N}}\\)\u003c/span\u003e\u003c/span\u003e to be less than one. Then, the DF part will bring negative weight for almost all words, probably except stop words. To avoid this issue, we add one to the result of division, before taking logarithm. Finally, by joining the modified IDF part, as DF, to TF part we can propose our weighting technique that is able to assign higher scores to more important news articles.\u003c/p\u003e \u003cp\u003eEquation \u003cspan refid=\"Equ5\" class=\"InternalRef\"\u003e5\u003c/span\u003e presents the proposed method of weighting features in financial news context. In this equation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({n}_{i}\\)\u003c/span\u003e\u003c/span\u003e is the number of documents containing word \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(i\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(N\\)\u003c/span\u003e\u003c/span\u003e denotes the total number of documents.\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$${W}_{i,j}= \\left(1+Log{f}_{i,j}\\right)\\text{*}Log (1+ \\frac{{n}_{i}}{\\text{N}} )$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Weighting Enhancement\u003c/h2\u003e \u003cp\u003eIn practical terms, textual analysis is a high dimensional data analysis in which each dimension represents a word. Term weighting scheme is a heuristic to calculate scores for dimensions. Finding the right score that can represent the true effect of dimension on the market is a complex problem. To overcome this problem, we follow up the term weighting adjustment with machine learning tactic.\u003c/p\u003e \u003cp\u003eThe problem we seek a solution for is about optimizing the outcome of term weighting technique. Therefore, we solve this problem using Genetic algorithm in a machine learning approach that is especially used for optimization problems.\u003c/p\u003e \u003cp\u003eLearning section in Genetic algorithm tries to optimize chromosomes, or in depth genes of a chromosome. Every gene of a chromosome is in a relationship with the weight of a term that can be explained by summation, multiplication or other simple or complex formula. According to empirical findings, we have settled on summation where we get better outcomes from adding term weight with gene.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e displays the implemented Genetic algorithm. After textual representation, each news article results in a feature vector. They in addition to the historical data and selected terms will be used to compute chromosomes fitness. If any of the end-conditions is satisfied in a generation, the iteration of Genetic algorithm will be stopped and the most fitted member in the last generation will be selected as the optimum solution.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eEvery chromosome is the subject of a fitness function and the result will be used as a tag of that chromosome. For as many as train set members (e.g., \u0026lsquo;n\u0026rsquo;) the process of fitness will run over every chromosome, hence there will be \u0026lsquo;n\u0026rsquo; tags associated with a chromosome, and its final solution calculated as the mean of \u0026lsquo;n\u0026rsquo; tags.\u003c/p\u003e \u003cp\u003eFitness value is calculated with absolute error. When fitness values are very close and the contrast between them needs to be magnified the choice will be squared error. The error is equal to distance of forecasted value and observed value. The average of these values will be taken to calculating general fitness via Eq.\u0026nbsp;\u003cspan refid=\"Equ6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. In this equation, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({A}_{i}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({F}_{i}\\)\u003c/span\u003e\u003c/span\u003e indicate observed and forecasted values in time of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(i\\)\u003c/span\u003e\u003c/span\u003e respectively, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e is the total number of train set members.\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$$General Fitness = \\frac{1}{n} \\sum _{i=1}^{n}\\left|{A}_{i}- {F}_{i}\\right|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThere are several common techniques for the selection. In this work, we employ an elitist method with negligible chance and the intention is to increase the chance of selecting elites for the next generation.\u003c/p\u003e \u003c/div\u003e"},{"header":"5 EVALUATION DESIGN","content":"\u003cp\u003eTo evaluate weighting techniques in the context of finance the optimum way is to assess them in a prediction. Information retrieval has its own evaluation practices to determine the efficacy of a weighting scheme but they are widely off-topic methods in financial analysis. We assess the competence of all models by performing a long-term forecasting on Tehran stock exchange overall index.\u003c/p\u003e \u003cp\u003eTehran stock exchange has been active since 1967. During this period, 672 companies have been engaged and it has grown into a large market that recently attracted researchers\u0026rsquo; attention ( [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e], [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e], [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]). There are 13 main and 45 industry indices on Tehran stock exchange where the overall index is the most significant and referred one among investors. Overall index measures the performance of 367 major components. The performance of the overall index in reality is deemed by traders as the performance of the stock market and then they always keep the index movement under careful observation in their investments.\u003c/p\u003e \u003cp\u003eA solid majority of companies that the overall index is under-effect their movements are governmental. Any qualitative data distributed from the government official media can affect their trends and hence the overall index. For example, news about any change in the exchange rate can bring about a huge effect on export-oriented companies; petroleum companies are under the effect of any change in the price of crack spreads or crude oil. All such and similar changes are issued and announced by the government itself. Many important governmental news is directly issued from the central bank media.\u003c/p\u003e \u003cp\u003eThe data set of this work includes financial news articles from the central bank of Iran and historical data from Tehran stock exchange from March 2005 until March of 2020. For evaluating the ultimate approach of this work, we utilized a function to cross validate test and train sets randomly and run it more than 10,000 times. From 178 months of collected data, 70% is chosen as train and the rest is used as test sets.\u003c/p\u003e \u003cp\u003eSeeking answer for research questions, we examine qualitative and quantitative observations. First, we delve into predicted trends. We analyze information that can be helpful for different trading strategies offered by heuristics in trend prediction. Next, we discuss quantitative evaluation results obtained from mean absolute error (MAE), mean absolute percentage error (MAPE), root mean squared error (RMSE) and directional accuracy (DA), which are some of the well-established assessing techniques and expressed in equations \u003cspan refid=\"Equ7\" class=\"InternalRef\"\u003e7\u003c/span\u003e to \u003cspan refid=\"Equ10\" class=\"InternalRef\"\u003e10\u003c/span\u003e respectively. In these equations, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({E}_{t}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({Y}_{t}\\)\u003c/span\u003e\u003c/span\u003e denote forecasting error and observed value at time \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(t\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(N\\)\u003c/span\u003e\u003c/span\u003e is the total number of testing, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(TR and TF\\)\u003c/span\u003e\u003c/span\u003e represent true directional forecasting and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(FR\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(FF\\)\u003c/span\u003e\u003c/span\u003e represent false directional forecasting.\u003cdiv id=\"Equ7\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$$MAE= \\frac{1}{N} {\\sum }_{t=1}^{N}\\left|{E}_{t}\\right|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ8\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e\n$$MAPE= \\frac{100}{N} {\\sum }_{t=1}^{N}\\left|\\frac{{E}_{t}}{{Y}_{t}}\\right|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e8\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ9\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ9\" name=\"EquationSource\"\u003e\n$$RMSE= \\sqrt{\\frac{1}{n}{\\sum }_{t=1}^{N}{E}_{t}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e9\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ10\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ10\" name=\"EquationSource\"\u003e\n$$Directional Accuracy= \\frac{TR+TF}{TR+TF+FR+FF}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e10\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eRMSE, as the root of MSE, and MSE both are useful to magnify the differences when errors are small or close together. On large errors, they generate very large numbers since they square the error. If a forecasting system gains small errors on average but rarely gains some large errors, then RMSE will be a very large number. Concretely, the general performance of a forecasting system cannot be fairly manifested by RMSE or MSE. In comparison, MAE and MAPE do not have bias on large numbers. Particularly MAPE that gauges the error relative to the actual value through dividing error by the observed value. Therefore, MAE and MAPE are preferable to determine the \u003cem\u003egeneral performance\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eMoreover, Useful information can be deduced from RMSE together with MAPE. For example, low MAPE appeared with high RMSE can attest to a performance that generally gains low error while cannot equally well handle large numbers.\u003c/p\u003e \u003cp\u003eIn similar fashion, directional accuracy together with closeness measures provide deductive information. A forecasting method with large closeness error will predict the direction with large margin from the actual value. It becomes misleading when it comes to predicting direction in a monotonous market in which direction changes only with tiny amounts. In such a situation, predicting trend with a high margin from actual value gives spurious information that the market is volatile; as well as it may fail to detect major trend movements. In comparison, a prediction with small margin betokens a stable market \u0026ndash; the information that those who are interested in buy and hold strategy are going after.\u003c/p\u003e"},{"header":"6 RESULTS AND DISCUSSION","content":"\u003cp\u003eGiven that we focus on evaluating the performance of heuristics, developing a competitive forecasting system is far from our scope. In our evaluation, forecasting is done merely using textual data, and results are pragmatic observations meant to assess and compare heuristics. They manifest the ability of heuristics in discovering predictive information. To answer the research questions, we delve into results and discuss experimental findings in terms of trend and error analysis. Magnified figures for discussions in section 6.1 are added in appendix.\u003c/p\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e6.1 Trend Analysis\u003c/h2\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e depicts observed and forecasted trends of Tehran stock exchange overall index. From 2005 until late 2009, the overall index is stable and consequently the forecasted values are relatively closer to the observed value. From late 2009 the trend gradually goes higher and then takes a sharp rise from mid-2012 to late 2013. From early 2014, it moves erratically until it starts a sharp rise in early 2018 again. Three considerable rising trends occur between the years of 2005 and 2020. First one starts in mid-2012 and moves the overall index from 23767 to 89723. The second and the smallest forms in late 2015 and lifts index from 61376 to 81723. This ends by the third stride that begins in early 2018, the index jumps from 92849 to 498700. In each specific period, there are possibilities in the form of risk and opportunity. We discuss supportive information that each model can offer.\u003c/p\u003e \u003cp\u003eFor the period 2005 to 2010, market moves almost on a horizontal line. Predicted trends by all models also attest to stability. In terms of predicting stability, the performance of all models are close. In late 2009 that the monotonous movement breaks, TF-IDF and TFDF methods are the first to notice that a nascent wave is emerging. When the market starts a more slope rise in mid-2012 it is revealed by TFDF, Probabilistic and Boolean, whereas TF-IDF detects a spurious oscillating movement. As a conclusion, in the period of 2005 to early 2014, all models perform closely in terms of predicting stability while in the matter of predicting incipient rising waves, which is useful for trend following strategy, TFDF is slightly better.\u003c/p\u003e \u003cp\u003eThere is fairly a lot error from early 2014 in all models however, predicted trends portend volatility that functions as a double-edged sword. Market instabilities are exaggerated by all models, highly by TF-IDF and slightly by TFDF. For those who are trading in short-terms with an eye toward technical analysis it signals a favorable opportunity that the market is swinging. They can benefit from recurrent rises and falls by making a swing-trading strategy. On the other hand, to those who are investing for long-term periods it indicates an inordinate amount of risk for buy and hold strategy. In such a situation, the market is not responding to fundamental data, which are the basis of buy and hold strategy. Therefore, they may better reconsider their investing decision or play it safe.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe wavering movement breaks in late 2015 and starts a sharp rise where the TF-IDF model places above the both predicted and observed trends. Technically, TF-IDF is forecasting a false downward movement. The TFDF model can predict and reveal the rising trend clearly. Following TFDF, probabilistic and Boolean models detect the trend, and TF-IDF is last to do. Again, in early 2016 market movement reverts to its oscillating fashion where all models perform almost the same. For the period of early 2014 to mid-2018, we settle that all models are able to notify volatility albeit in different scales. In addition, the emerging ascend is clearly predicted by TFDF at the right time where other models do with a fair delay.\u003c/p\u003e \u003cp\u003eAfter a lot of bounce for almost five years, since mid-2018, trend starts a big, steep rise. Remarkably, the TFDF model can predict the rise right before the starting time window of the sharp rise where TF-IDF fails and predicts a misleading fall. The Boolean and Probabilistic models are also able to reveal the start line of the trend though they fail to continue the trend. For the next rising trend, the first to detect movement is TF-IDF followed by TFDF. For the last trend, TFDF performs slightly better than other models. As a result, for the period of mid-2018 to early 2020, TFDF is able to predict all trends. Moreover, on the matter of time, TFDF reveals the major inchoate trend at the right time similar to Boolean and Probabilistic but more closely. TF-IDF shows weakness in predicting the rising trend at the propitious moment.\u003c/p\u003e \u003cp\u003eAs a conclusion, in predicting rising trends, the TFDF model can present a sound support for trend following decisions. In the matter of revealing volatility, we find that all models are able to do well but they predict fluctuations in different scales where traditional models exaggerate volatility more than our model. Most of the time, we see TF-IDF delivers a wavering trend, a signal which is translated to volatility, that is not a true barometer of actual trend. The information revealed by TFDF can boost technical based decisions in weekly and monthly investments. As well as it can warn fundamentalists about an unresponsive market to the intrinsic values, similar to the other models and on a more precise scale.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e6.2 Error Analysis\u003c/h2\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e illustrates the absolute error of models in a logarithmic scale. Markers are representing the forecasting points. The amount of errors is depicted by how high the markers are above from the bottom. Therefore, the lower error leads to the lower place. We use logarithmic scale to avoid the high contrast map of errors that occurs in arithmetic scale. In the arithmetic scale, small errors and changes are shown in an imperceptible degree while large numbers are exceedingly magnified.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe probabilistic model is placed above other models almost all the time and forms a line of circle markers on top. For most parts, it is like a competition for taking the lower places between Probabilistic and Boolean and for some parts, it occurs between TF-IDF and Probabilistic especially after mid-2018. It is even observable previously in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e that the TF-IDF trend shows a higher buoyancy in that period. Surprisingly, the lowest points are registered by TF-IDF and Boolean models. From looking at the whole, the cloud of markers is demarcated by TFDF markers from the bottom, except some points. The line is not as solid as the one formed by the Probabilistic model. In some parts, there are competitions between most of the models to take lower places whereas for the most parts it is going on between TFDF and TF-IDF models. In conclusion, for most of time the figure shows that TF-IDF takes lower places than other traditional models, while it delineates a consistent tendency for TFDF to take lower places among all models.\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e displays the closeness of the predicted values. Smaller numbers in this table indicate more precision. The probabilistic model has the worst results in all measures and with a small amount Boolean model places in the next. While in RMSE, TF-IDF gains slightly higher error than Boolean, it reduces MAPE by 48 percent. RMSE is used to magnify the error to make sharp contrast in comparison. On the other side, MAPE is used to disclose the general error by calculating error relative to the actual value. In addition, prediction on large numbers is indeed conducive to large errors. Therefore, higher error in RMSE and lower error in MAPE can demonstrate that Boolean generates more error in general than TF-IDF. On the other hand, TF-IDF creates more error when it comes to predicting large values. Yet, it demonstrates that TF-IDF is much more reliable in general, relative to Probabilistic and Boolean. The lower absolute error in MAE by TF-IDF substantiates the suggestion. It may uphold the use of TF-IDF as the primary weighting method in textual analysis.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMean of errors\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod / Measures\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMAE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMAPE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBoolean\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15418\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e11163\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e27.29\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProbabilistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e18099\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e12939\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e31.06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTF-IDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15739\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9763\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e18.35\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTFDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e10455\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7109\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e15.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eOur model produces result of unmatchable precision in all measures. The lowest RMSE among traditional models is obtained by Boolean and TFDF lowers it by 47 percent. In terms of absolute error, the lowest MAE is obtained by TF-IDF then TFDF reduces it by 37 percent. Again, TF-IDF has the lowest MAPE relative to Boolean and Probabilistic but TFDF performs superiorly and achieves 17 percent more precision by reducing MAPE from 18.35 to 15.59.\u003c/p\u003e \u003cp\u003eDirectional accuracy is summarized in Table \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. Theoretically, it is possible to guess the direction of movement 50 percent true only by chance. However, it may not be achievable randomly in the real world. Why so, the occurrence of rise and fall has never been evenly distributed. In practice, trends have bias and in the case of overall index, we observe the buoyancy of the Tehran stock exchange. Therefore, the thought of achieving 50 percent by chance is not a just consideration.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDirectional accuracy\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod / Measure\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDirectional Accuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBoolean\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e52.81\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProbabilistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e51.12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTF-IDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e56.74\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTFDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e57.30\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eBy the first glance at the table, we see that the Probabilistic model acquires the least accuracy, as the same performance in closeness. TF-IDF performs the best among traditional models as the same act in absolute error. The Boolean model is not far ahead of probabilistic where it achieves 3 percent higher accuracy. TFDF presents greater results among all. Although it precedes TF-IDF only by a nominal margin, reliable performance in trend prediction reinforces its result. In contrast, poor performance in predicting trend blunts good result of TF-IDF.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e6.3 Answer to Research Questions\u003c/h2\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e reviews precision and accuracy results of all models relative to TF-IDF. Bars represent errors so the lower is the better, except in directional accuracy (DA). Relative performances shown in the chart are useful to analyze how a model acts compared with others. Pursuing answer for the first research question, we scale results relative to TF-IDF that makes its performance clearer. For example, it shows great performance of TF-IDF in MAPE where Probabilistic and Boolean gain 69 and 48 percent more error respectively.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo answer the first research question, we investigate and discuss the forecasting results. \u003cem\u003eFirst\u003c/em\u003e, we set the actual moves of the overall index and predicted trends side by side, from 2005 to 2020. In terms of forecasting trend, TF-IDF is able to reveal volatility; in a few cases detects rising trend; sometimes gives spurious signals; and for the most part its trend is oscillating. Consequently, the inconstant performance renders the predicted trend invalid. \u003cem\u003eSecond\u003c/em\u003e, we examine closeness measures data of all models. On the matter of closeness, which lower is better, we assess models using three measures. RMSE data indicate that TF-IDF is far better than Probabilistic although with a slight margin it falls behind the Boolean. We observe vivid error reduction in MAE where TF-IDF wins against other information retrieval models, by generating the lowest absolute error. TF-IDF achieves an unbeatable score in MAPE where it was able to decrease the lowest error dramatically, which was from Boolean, by 48 percent. \u003cem\u003eThird\u003c/em\u003e, we assess the accuracy of predicting direction for which we gather data from monthly forecasting. In that measure, TF-IDF gets itself in front but it is unreliable due to wavering trend prediction. In conclusion, except for the weakness in trend prediction, TF-IDF evidences better performance in general and is the best among major information retrieval models.\u003c/p\u003e \u003cp\u003eComparative results of TFDF and TF-IDF are summarized in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e relative to TFDF. To make improvements clearer we scale results relative to the proposed model. It is even obvious to the casual observer that how unproductive is TF-IDF compared with TFDF, in closeness measures. The most improvement occurs in RMSE and the least does in directional forecasting. Given that, high RMSE can betoken of weakness in handling large numbers, it shows that the proposed model improves performance on large numbers by 50 percent.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo answer the second research question, we field our model in a truly equal race among all models. \u003cem\u003eFirst\u003c/em\u003e, we juxtapose the observed and forecasted trends of the overall index by all models. We consider information that is useful for short-term and long-term investments. On that matter, we analyze the information proffered by all models for different trading strategies. TFDF is greatly able to detect rising trends, which is useful for technicians especially those who are interested in trend-following strategy. It is the first to reveal all major rising trends in their nascent state. For those who are trading in short-terms TFDF is able to signal the opportune moments when volatility is escalating. As well as for fundamentalists it is able to show when the market is responsive to intrinsic values and when the market is tempestuous and too risky for buy and hold strategy. The proposed model performs better than any other did in trend forecasting. \u003cem\u003eSecond\u003c/em\u003e, in forecasting values we gauge how close the predicted values are to the real values. We examine the distances between actual and predicted values and observe that TFDF provides most of the shorter distances. On average, TFDF reduces absolute error by 37 percent. In other measures such as RMSE, our model gains the lowest error and increases closeness by 50 percent. Likewise, the proposed model decreases TF-IDF error in MAPE, which was the finest, by 17 percent. In closeness measures, TFDF produces unsurpassed results and reduces error by 34.6 percent on average. \u003cem\u003eThird\u003c/em\u003e, we compare directional accuracy by examining the predicted directions in all 178 months and find that TFDF delivers the highest accuracy. There is no big difference between TFDF and TF-IDF in directional accuracy raw data but reliable performance in trend prediction reinforces such data. As a conclusion, the TFDF model benefits from our heuristic and evidently supports our hypothesis that news stories with higher publication rate provide more predictive information.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e6.4 Enhanced Model and Closing Notes\u003c/h2\u003e \u003cp\u003eIn Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, performance of TFDF is matched against when it is aided by machine learning optimization, in terms of closeness measures. Machine learning enhances TFDF in all ways. In RMSE, where we see the lowest enhancement, machine learning lowers error by 14.32 percent. Then, it reduces the amount of absolute error by 15.70 percent. The most enhancement occurs in MAPE measure, which shows 19.37 percent more precision. The lowest error reduction obtains in RMSE and the greater in MAPE that forms a pattern similar to TF-IDF performance. Previously observed, TF-IDF is beaten by Boolean in RMSE with a small margin while acts excellently in MAPE and shrinks Boolean error by 48 percent. It is sensible to say that Genetic optimization suffers from the same difficulty as we mentioned for TF-IDF in forecasting large numbers, though it happens to a very smaller degree.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMean of errors\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod / Measures\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMAE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMAPE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTFDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e10455\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7109\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e15.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEnhanced TFDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e9145\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6144\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e13.06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e compares directional accuracy of both TFDF and the enhanced model with machine learning. Here we encounter unexpected results when TFDF is enhanced via Genetic algorithm. Despite the palpable error reduction in closeness measures, in directional forecasting the TFDF model is ahead, although by a meagre amount. We expect more accuracy as we get more precision but it does not happen nonetheless.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDirectional accuracy\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod / Measure\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDirectional Accuracy\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTFDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e57.30\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEnhanced TFDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e57.14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eWe have an assumption. When variables of a prediction are close and trend moves implicitly, directional forecasting becomes more challenging than when the market is explicitly rising or falling. We cross validate test and train sets randomly and run the test more than 10,000 times. It may accidentally happen that test sets are filled, more than its share, with times from periods when market has been in the doldrums that eventually costs the trivial loss in directional accuracy.\u003c/p\u003e \u003cp\u003ePrecision and accuracy results of TFDF are summarized and compared with results when weights are enhanced using machine learning in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e. To make obtained achievements by machine learning clearer we scale results relative to the enhanced model. Except for directional forecasting, machine learning succeeds in improving performance. The most change occurs in MAPE that can attest to general improvement in precision.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eOther studies train machine often to help predict trend movement. They develop classification systems mainly with the help of support vector machine. We train machine to optimize the weight of words. Our approach delivers perceptible effects on precision. On average, it helps to achieve 16.46 percent more precision in predicting value. Unexpectedly, it faces a trivial loss in directional accuracy compared with TFDF, while it is still better than other models. There is a reasonable prospect that directional accuracy improves with the help of a classifier, which may be considered as a minor future work.\u003c/p\u003e \u003cp\u003eSchumaker and Chen argue that the weakness of bag of words stems from a vast collection of various words that works as a real handicap (bag of words is analogous to the TF-IDF model in our work). They suggest the better performance by other models comes from truncated representation of documents. However, it is not a cogent argument in our work regarding two reasons. \u003cem\u003eFirst\u003c/em\u003e, we select the most influential words from the corpus that precludes bag of words from suffering a vast collection of words. \u003cem\u003eSecond\u003c/em\u003e, all models use the same word list that makes the measurement equal in all common ways. Therefore, there is no noisy word list nor a real handicap that gives reasons for the weakness in bag words; rather the weakness is inherent in its procedures where our findings reveal that the deficiency lies in the course of weighting terms.\u003c/p\u003e \u003cp\u003eIn brief, all heuristics unveil their ability in search of predictive information. The Probabilistic model produces poorest results in directional accuracy also in closeness measures. Boolean acts better than TF-IDF in predicting trend and gains lowest RMSE among information retrieval models. TF-IDF brings the finest results among major information retrieval models. In MAPE and MAE it gains the lowest error but falters in trend prediction. It answers our first research question and can justify why TF-IDF has been reckoned the standard model in textual analysis. However, none of the borrowed heuristics from information retrieval shows satisfying performance in all measures that brings up the proposition: the finance area should consider study on text-processing techniques as one of its own pursuits. On that matter, we propose a financial related heuristic. Our model outmaneuvers all models in all measures. The TFDF achieves highest precision and directional accuracy. It is the only model that detects all major rising trends in their nascent state. According to the general rules of forecasting, all models conduct themselves reasonably; however, the superiority of TFDF demonstrates that our heuristic is more consistent with financial forecasting bylaws. It answers our second research question and confirms that news with higher publication rate provides more predictive information.\u003c/p\u003e \u003c/div\u003e"},{"header":"7 CONCLUSION AND FUTURE WORKS","content":"\u003cp\u003eAmong major information retrieval models, TF-IDF is known as the de facto standard scheme in textual analysis. Its heuristic is perfectly compatible with problems in its origin topic. As the first issue considered in this work, we examine major information retrieval models in order to find whether TF-IDF is efficient in resolving business issues and if it is better than other standard models.\u003c/p\u003e \u003cp\u003eWe find that information retrieval models do not afford forecasting system the ability to detect all major movements. TF-IDF often delivers belated signs of rising trends and falls behind Boolean in RMSE by a small margin. In other closeness measures, it surpasses all standard models by providing 31.53 percent more precision on average. In directional accuracy, the Probabilistic model has the poorest result and though TF-IDF is ahead of Boolean by 6.93 percent, its result is vitiated by wavering trend prediction. Consequently, Boolean is preferable in trend prediction and forecasting on large numbers whereas TF-IDF acts superiorly in general according to MAPE and MAE measures.\u003c/p\u003e \u003cp\u003eText-processing techniques have not been improved or modified in compliance with norms in the context of business. They are merely borrowed from other areas, chiefly from information retrieval. We believe that nonbusiness text-processing techniques are incompatible with financial forecasting and potential to bring unproductive weights in the course of scoring words. It motivates us to work on a new weighting heuristic especially devised for financial analysis, as the main issue we consider in this work.\u003c/p\u003e \u003cp\u003eWe develop our new weighting heuristic upon the hypothesis: news stories with higher publication rate contain more predictive information. We diagnose the TF-IDF heuristic and find the problematic part, then we modify it to calculate a higher score for more important information. We maintain that the \u0026ldquo;inversion of document frequency\u0026rdquo; tends to vitiate the term weighting process. The acute shortcoming becomes manifest in comparing the results of TF-IDF with our model as TFDF.\u003c/p\u003e \u003cp\u003eThe proposed model shows reliable performance in discovering rising trends where TF-IDF fails. In closeness measures, it reduces errors of TF-IDF by 38.5 percent on average. TF-IDF provides competitive performance in directional accuracy but it is still beaten by TFDF nonetheless. Overall, the proposed heuristic brings better performance in all measures and evidently supports our hypothesis.\u003c/p\u003e \u003cp\u003eIn additional analysis, to overcome the high dimensionality of textual analysis we suggest optimizing scores using Genetic algorithm. Although a trivial loss surfaces in directional accuracy, it does not cast a shadow on the efficiency of the approach. We achieve 19.37 percent improvement in MAPE whereas 15.70 and 14.32 percent progress in MAE and RMSE respectively. On average, the enhanced method reduces the closeness errors by 16.46 percent.\u003c/p\u003e \u003cp\u003eOur findings show that information retrieval solutions are not meant to succeed in the realm of business. Our model is not infallible but its performance still discloses how abortive information retrieval models are in discovering predictive information. They are designed for retrieval purposes that narrows their ideal efficacy within the retrieval\u0026rsquo;s orbit.\u003c/p\u003e \u003cp\u003eOur evaluation focuses on revealing models efficiency in term weighting. Pursuing that, we use a word list in order to bypass the issues consequent upon feature selection and to ensure that models are not encumbered with any unforeseen disadvantage in preprocessing procedures. Developing word list manually is not a time-effective approach so in practical terms especially when data are large and mounting up regularly it may be recommended to select features automatically. As a future work, we suggest running feature selection using automatic settings in order to examine how models perform this section, thereupon the entire analysis as well.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthical Approval\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eNot applicable\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eNot applicable\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAll authors have contributed to the writing and reviewing of the manuscript text and the level of contribution is according to the name list order maintained in the paper.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research and publication was not supported by any grant.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAll the materials including codes, experimental road map, data collected are available and ready to be provided as requested by the reviewers.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eW. Antweiler and M. Z. Frank, \u0026quot;Is All That Talk Just Noise? The Information Content of Internet Stock Message Boards,\u0026quot; The Journal of finance, vol. 59, no. 3, pp. 1259-1294, 2004.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eR. P. Schumaker and H. Chen, \u0026quot;Textual Analysis of Stock Market Prediction Using Breaking Financial News: The AZFinText System,\u0026quot; ACM Transactions on Information Systems (TOIS), vol. 27, no. 2, pp. 1-19, 2009.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eM. Arias, A. Arratia and R. Xuriguera, \u0026quot;Forecasting with twitter data,\u0026quot; ACM Transactions on Intelligent Systems and Technology (TIST), vol. 5, no. 1, pp. 1-24, 2014.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eJ. Bollen, H. Mao and X. Zeng, \u0026quot;Twitter mood predicts the stock market,\u0026quot; Journal of Computational Science, vol. 2, no. 1, pp. 1-8, 2011.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eH. N. Ozsoylev, JohanWalden, M. D. Yavuz and R. Bildik, \u0026quot;Investor Networks in the Stock Market,\u0026quot; The Review of Financial Studies, vol. 27, no. 5, pp. 1323-1366, 2014.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eE. F. Fama, The behavior of stock-market prices, vol. 38, no. 1, pp. 34-105, 1965.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eB. G. Malkiel, A Random Walk Down Wall Street, New York: WW Norton \u0026amp; Company, 1999.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eR. Baeza-Yates and B. Ribeiro-Neto, Modern information retrieval, New York: ACM press, 1999.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eT. Loughran and B. McDonald, \u0026quot;When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks,\u0026quot; The Journal of finance, vol. 66, no. 1, pp. 35-65, 2011.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eH. N. Semiromi, S. Lessmann and W. Peters, \u0026quot;News will tell: Forecasting foreign exchange rates based on news story events in the economy calendar,\u0026quot; The North American Journal of Economics and Finance, vol. 52, p. 101181, 2020.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eM. Hahsemi, M. Rezaei and M. Kaedi, \u0026quot;Textual analysis of central bank news in forecasting long-term trend of Tehran stock exchange index,\u0026quot; Journal of Information and Communication Technology, vol. 12, no. 44, pp. 119-132, 2020.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eM. Butler and V. Ke\u0026scaron;elj, \u0026quot;Financial forecasting using character n-gram analysis and readability scores of annual reports,\u0026quot; in Canadian Conference on Artificial Intelligence, Berlin, Heidelberg, 2009.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eM.-A. Mittermayer, \u0026quot;Forecasting intraday stock price trends with text mining techniques,\u0026quot; in Proceedings of the 37th Hawaii International Conference on Social Systems, Big Island, HI, USA, 2004.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eY.-W. Seo, J. A. Giampapa and K. Sycara, \u0026quot;Text Classification for Intelligent Portfolio Management,\u0026quot; Robotics Institute, Carnegie Mellon University , Pittsburgh, PA, 2002.\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eJ. Thomas and K. Sycara, \u0026quot;Integrating genetic algorithms and text learning for financial prediction,\u0026quot; in Genetic and Evolutionary Computation Conference (GECCO), 2000.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eG. Gidofalvi and C. Elkan, \u0026quot;Using news articles to predict stock price movements,\u0026quot; Department of Computer Science and Engineering, University of California, San Diego, 2001.\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eM. R. Amin-Naseri and E. A. Gharacheh, \u0026quot;A hybrid artificial intelligence approach to monthly forecasting of crude oil price time series,\u0026quot; in 10th International Conference on Engineering Applications of Neural Networks, 2007.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eS. Feuerriegel and J. Gordon, \u0026quot;Long-term stock index forecasting based on text mining of regulatory disclosures,\u0026quot; Decision Support Systems, vol. 112, pp. 88-97, 2018.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eA. Sitaram and B. A. Huberman, \u0026quot;Predicting the future with social media,\u0026quot; in IEEE/WIC/ACM international conference on web intelligence and intelligent agent technology, Toronto, ON, Canada , 2010.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eA. Tumasjan, T. Sprenger, P. Sandner and I. Welpe, \u0026quot;Predicting elections with twitter: What 140 characters reveal about political sentiment,\u0026quot; in Fourth International Conference on Weblogs and Social Media, Washington, DC, USA, 2010.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eA. Culotta, \u0026quot;Detecting influenza outbreaks by analyzing Twitter messages,\u0026quot; 2010.\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eV. Lampos, T. D. Bie and N. Cristianini, \u0026quot;Flu detector-tracking epidemics on Twitter,\u0026quot; in Joint European conference on machine learning and knowledge discovery in databases, Berlin, Heidelberg, 2010.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eK.-G. Aase, \u0026quot;Text Mining of News Articles for Stock,\u0026quot; Institutt for datateknikk og informasjonsvitenskap , Trondheim, Norway, 2011.\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eB. O\u0026apos;Connor, R. Balasubramanyan, B. Routledge and N. Smith, \u0026quot;From tweets to polls: Linking text sentiment to public opinion time series,\u0026quot; in Proceedings of the International AAAI Conference on Weblogs and Social Media, Washington, DC, 2010.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eG. G.-R. Wu, T. C.-T. Hou and J.-L. Lin, \u0026quot;Can economic news predict Taiwan stock market returns?,\u0026quot; Asia Pacific management review, vol. 24, no. 1, pp. 54-59, 2019.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eA. S. A. Rahman and S. Abdul-Rahman, \u0026quot;Mining textual terms for stock market prediction analysis using financial news,\u0026quot; in International Conference on Soft Computing in Data Science, Singapore, 2017.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eW.-B. Yu, B.-R. Lea and B. Guruswamy, \u0026quot;A Theoretic Framework Integrating Text Mining and Energy Demand Forecasting,\u0026quot; IJEBM, vol. 5, no. 3, pp. 211-224, 2007.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eJ. M. Keynes, The General Theory of Employment, Interest, and Money, Harcourt, 1936.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eP. Falinouss, \u0026quot;Stock trend prediction using news articles: a text mining approach,\u0026quot; Lulea University of Technology, 2007.\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eM. Fasanghari and G. A. Montazer, \u0026quot;Design and implementation of fuzzy expert system for Tehran Stock Exchange portfolio recommendation,\u0026quot; Expert Systems with Applications, vol. 37, no. 9, pp. 6138-6147, 2010.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eJ. Zahedi and M. M. Rounaghi, \u0026quot;Application of artificial neural network models and principal component analysis method in predicting stock prices on Tehran Stock Exchange,\u0026quot; Physica A: Statistical Mechanics and its Applications, vol. 438, pp. 178-187, 2015.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eP. Shahrestani and M. Rafei, \u0026quot;The impact of oil price shocks on Tehran Stock Exchange returns: Application of the Markov switching vector autoregressive models,\u0026quot; Resources Policy, vol. 65, p. 101579, 2020.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eR. Ramezanian, A. Peymanfar and S. B. Ebrahimi, \u0026quot;An integrated framework of genetic network programming and multi-layer perceptron neural network for prediction of daily stock return: An application in Tehran stock exchange market,\u0026quot; Applied soft computing, vol. 82, p. 105551, 2019.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cspan style=\"white-space:pre;\"\u003e\u0026nbsp; \u0026nbsp;\u0026nbsp;\u003c/span\u003eA. H. Ghahfarrokhi and M. Shamsfard, \u0026quot;Tehran stock exchange prediction using sentiment analysis of online textual opinions,\u0026quot; Intelligent Systems in Accounting, Finance and Management, vol. 27, no. 1, pp. 22-37, 2020.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Financial textual analysis, Term weighting, Machine learning, Stock market ","lastPublishedDoi":"10.21203/rs.3.rs-2883673/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2883673/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eTextual analysis in the realm of business depends on text processing techniques borrowed mainly from information retrieval. However, it is not a viable solution capable of developing in finance. We suggest developing financial homegrown techniques for processing textual data. Especially in the course of scoring words where standard techniques are incongruous in financial analysis. On that matter, we pursue two issues. \u003cem\u003eFirst\u003c/em\u003e, we examine major information retrieval heuristics. We find TF-IDF a facile solution that falters in predicting trend and generates high errors on large numbers. \u003cem\u003eSecond\u003c/em\u003e, we work on a new heuristic satisfying financial concernments. We consider the relation between publication rate of information and their importance. The proposed heuristic provides results of unmatchable performance in both predicting trend and precision measures. In additional analysis, we optimize our scheme using Genetic algorithm in a machine learning approach and get greater precision. In comparison with TF-IDF, the proposed heuristic conduces to 38.5 percent lower error in closeness measures which is again reduced by 16.46 percent with the help of machine learning. Our findings suggest that financial textual analysis and information retrieval should go their separate ways.\u003c/p\u003e","manuscriptTitle":"TFDF and TF-IDF in Financial Analysis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-05-05 21:53:50","doi":"10.21203/rs.3.rs-2883673/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"3140eb6c-efb2-49ca-8385-5ecfeffe0d3b","owner":[],"postedDate":"May 5th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-05-09T08:44:34+00:00","versionOfRecord":[],"versionCreatedAt":"2023-05-05 21:53:50","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-2883673","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2883673","identity":"rs-2883673","version":["v1"]},"buildId":"cBFmMYwuxLRRLfASyISRj","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0