A BERT-Based Framework for Extracting Business Insights from Financial Reports

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Financial reports bear a lot of textual information, which is normally unstructured and imperative to business decision-making. Nevertheless, it is difficult to draw meaningful conclusions out of such huge and sophisticated papers through the conventional methods of analysis. This paper is an attempt to provide a BERT-based system to extract business insights in financial reports in an automated manner. Using the ability of Bidirectional Encoder Representations of Transformer to contextualize textual information (annual reports, management discussions, financial statements, etc.), the model proposed then analyzes the textual information to extract important information (sentiment, risk factors, performance indications, etc.). The framework includes the data preprocessing, domain-specific weighting of a pre-trained BERT, and usage of classification and information extracting methods. The research experimental results prove that BERT-based system is much more operative than the old-style machine learning (ML) models in terms of accuracy, precision and F1-score. Moreover, the inferences that were to be made are of good assistance to all stakeholders- investors, managers and analysts for making sound decisions. The paper is relevant to the area of financial text analytics research because it provides a scalable and efficient framework in processing unstructured financial information. Future development can be based on real-time analytics and connection to financial decision support systems.
Full text 127,149 characters · extracted from preprint-html · click to expand
A BERT-Based Framework for Extracting Business Insights from Financial Reports | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A BERT-Based Framework for Extracting Business Insights from Financial Reports Prasant Kumar Rout, Binita Nanda, Sunil Kumar Dhal This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9209699/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Financial reports bear a lot of textual information, which is normally unstructured and imperative to business decision-making. Nevertheless, it is difficult to draw meaningful conclusions out of such huge and sophisticated papers through the conventional methods of analysis. This paper is an attempt to provide a BERT-based system to extract business insights in financial reports in an automated manner. Using the ability of Bidirectional Encoder Representations of Transformer to contextualize textual information (annual reports, management discussions, financial statements, etc.), the model proposed then analyzes the textual information to extract important information (sentiment, risk factors, performance indications, etc.). The framework includes the data preprocessing, domain-specific weighting of a pre-trained BERT, and usage of classification and information extracting methods. The research experimental results prove that BERT-based system is much more operative than the old-style machine learning (ML) models in terms of accuracy, precision and F1-score. Moreover, the inferences that were to be made are of good assistance to all stakeholders- investors, managers and analysts for making sound decisions. The paper is relevant to the area of financial text analytics research because it provides a scalable and efficient framework in processing unstructured financial information. Future development can be based on real-time analytics and connection to financial decision support systems. BERT Financial Report Analysis NLP Sentiment Analysis Risk Factor Extraction FinBERT Business Intelligence Transformer Models Financial Text Analytics Figures Figure 1 1. Introduction 1.1 Importance of Financial Reports in Business Decisions The basis of corporate transparency and communication with stakeholders promote financial reporting. Annual reports, quarterly filings, management discussion and analysis (MD&A) applications and earnings call transcripts are all the most complete and authoritative data of the performance of a company, objectives, and second outlook. The documents will be the foremost foundation of information to a variety of parties, such as equity investors to review their positions in the portfolio, credit analysts to analyze the risks of default, regulatory authorities to ensure compliance with their regulation, and managers of the corporation to compare themselves with others in the respective industries. The strategic worth of financial reports is far more than the arranged numerical tables that they consist of. Narrative disclosures such as management commentary relating to competitive positioning, descriptions of risk factors and future guidance are likely to hold the wealthiest and most useful intelligence, even though they are no longer accessible to automated intelligence-gathering systems. Those organizations that are able to quickly and properly interpret these stories are able to have a decisive information advantage in the financial markets. 1.2 Challenges of Analyzing Unstructured Text Financial reports in their current form are daunting with respect to automated analysis in spite of their informational richness. An average SEC 10-K filing will be between 100 and 300 pages of formatted balance sheets and income statements which is filled with thick legal wading explanatory text. There are a number of features that make this text particularly hard to analyze: Volume and complexity: A single annual report may contain over 100,000 words across dozens of sections with highly specialized vocabulary Domain specificity: Financial language employs technical terminology ("amortization", "goodwill impairment", "covenant breach") whose meaning differs substantially from everyday usage Contextual ambiguity: Identical phrases carry opposite sentiment in different contexts - "above average" can be positive (performance) or negative (risk) Document heterogeneity: Reports vary significantly in structure, length, and writing style across companies, industries, and reporting periods Implicit information: Critical insights are often embedded in nuanced language, hedging constructions, and comparative references that require deep comprehension to interpret 1.3 Rise of AI and NLP in Finance With the intersections of progress in Natural Language Processing (NLP), and the proliferation of more digitally-published financial materials, a novel challenge of automated financial text analysis has become an opportunity of an unprecedented scale. Use of AI systems in the financial industry is being actively implemented in applications such as textual sentiment analysis of trading information, compliance filings, forecasting earnings surprises, and credit risk estimation by textual disclosure as well as news sentiment analysis to give trading signals. Transformer revolution - launched by Vaswani et al. (2017) [ 13 ] and quickly followed by the likes of BERT, GPT and RoBERTa - has dramatically increased the ability of NLP systems to reason about complex, context unique language. Naturally trained networks, which learn representations based on billions of words of text, can transfer well to the domain of particular specialized task, with relatively little fine-tuning. 1.4 Limitations of Traditional Methods Conventional techniques used to analyze financial text have been limited to rule-based and statistical techniques which are poorly adapted to the complexity of the financial language of today: Lexicon-based methods (e.g., Loughran-McDonald dictionary): Context-blind, static, cannot adapt to new terminology or negations Bag-of-words models with TF-IDF: Ignore word order and syntactic structure; suffer from high dimensionality and sparsity Support Vector Machines and Naive Bayes: Dependent on handcrafted features; poor generalization across document types Early neural networks (LSTM, RNN): Capture sequential structure but struggle with long-range dependencies and require large labeled datasets These limitations result in systematic failures to capture the contextual, relational, and domain-specific dimensions of financial language that are most critical for insight extraction. 1.5 Problem Statement Paper-based and other conventional automated approaches do not find rich and context-driven insights in the complex financial text. This is because of the absence of domain-specific knowledge, long document processing inability, and the inability to deal with subtle financial language leading to incomplete, inaccurate, and non-scalable financial intelligence. It is of paramount importance to have a smart, automated system that would be reliable in deriving actionable business information out of scale unstructured financial reporting 1.6 Objectives The following are the main objectives to be used in this study: Create a Bert-based NLP system adapted to financial report analysis extraction- extract validation of major business insights on unstructured financial text, such as sentiment signals, risk factors and performance indicators Compare the action of the proposed BERT-based system to the known baseline systems such as lexicon based models, conventional machine learning and previous deep-learning models Discuss practical usefulness of extracted insights to the financial stakeholders such as investors, analysts and managers 2. Literature Review This section is a review of the history of research of the early financial text analytics through the traditional methods of machine learning, deep learning methods, and finally transformer- based models where the proposed framework was built. 2.1 Financial Text Analytics Computational finance of financial text has a long history that originates decades earlier. Initial studies acknowledged that predictive signatures were contained in textual information of financial reports and were not completely represented by numerical data. It is based on this foundational relationship amongst the tone of texts and stock market performance that the negative language present in Wall Street Journal columns is predictive of stock markets declines (Tetlock, 2007) [ 12 ]. Later endeavors were extended to the area of corporate disclosures. Loughran and McDonald (2011) [ 7 ] developed an important input by building the first lexicon of finance specific sentiment lexicon showing that negative dictionary words (using general purpose dictionary, e.g. Harvard General Inquirer) have neutral or technical senses in financial life. Their lexicon of words, such as positive, adverse, ambiguity, litigious and restricting words, was adopted as the standard of the financial NLP research almost ten years. Since then, NLP studies of finances have been developed to include various types of documents, such as quarterly earnings call transcripts (Qin and Yang, 2019) [ 10 ], analyst reports, or regulatory documents as well as social media options. These have usage in return prediction and volatility forecasting to credit risk assessment and fraud detection. 2.2 Traditional Approaches 2.2.1 Lexicon-Based Approaches Lexicon built approaches provide sentiment scores to documents by matching words in documents with predefined word lists and aggregating those scores at the sentence/document level. The most widely used lexicon that is cited in finance to gauge sentiment is Loughran- McDonald (LM) classical dictionary [ 7 ], the advanced dictionaries have since been created to process the language of the earnings calls, credit ratings statements, and central bank statements. Whereas lexicon based methods do create transparency, interpretability and are efficient in terms of computation, limitations are well documented. They are contextually blind in their essence, i.e. "outstanding" is a positive term whether they are speaking of an outstanding performance or outstanding liabilities. The aspects of negation (not profitable, no significant improvement), are either missing or restricted to mere negation windows. Premeditated vocabularies are not able to adjust to the changes of financial terms, changes in regulations language, and jargon usage. 2.2.2 Machine Learning Approaches (SVM, Naive Bayes) Supervised machine learning models also overcame certain limitations of lexicon based methods by having classification limits learned over labelled training data and not using a set of predefined rules. TF-IDF features upon Support Vector Machines (SVMs) in the classification of financial text proved a formidable baseline in financial text classification tasks, and was proven competitive against sentiment classification and topic categorization benchmarks. Naive Bayes classifiers were based on the assumption of conditional independence and were computationally efficient and strong to small datasets, so they were used in the classification of financial news. Ensemble models such as the Random Forests and Gradient Boosting models have also been shown to achieve performance that is even better in structured sets of features that are based on both textual and non-textual cues. Nevertheless, these gains could not overcome the fact that the conventional machine learning models were limited in terms of requiring hand-crafted features, bag-of-words models that ignored the order of words in a document, and generalization of their applications across financial sectors and documents. 2.3 Deep Learning Approaches 2.3.1 Recurrent Neural Networks (RNN) Sequential Processing NLP Recurrent Neural Networks added the concept of state representations to be hidden along position of tokens in models. On financial text, RNNs also learnt longitudinal dependencies, made them become more accurately captured than bag-of-words models. Nevertheless, vanilla RNNs had problem of vanishing gradient which inhibited their capability to detect long-range dependencies essential in the interpretation of complex financial sentences. 2.3.2 Long Short-Term Remembrance Networks (LSTM) Long Short-Term Remembrance networks (Hochreiter & Schmidhuber, 1997) addressed the vanishing incline problem through fenced memory cells that selectively retain and reject information across long sequences. Bidirectional LSTMs (BiLSTMs) further extended this by processing systems in both forward and backward directions, providing richer contextual representations. In earnings call transcript prediction, Qin and Yang (2019) [ 10 ] used LSTMs to predict stocks, and reported higher performance than lexicon-based baselines. As mechanisms of attention, models that are paired with LSTMs (BiLSTM-Attention) enabled them to dynamically assign different weights to contributions of the various positions of the sentence, enabling these models to enhance their financial classification performance. Nevertheless, LSTMs still had some weaknesses such as their sequential processing limitation, ineffectiveness in being able to parallelize training and still had problems with very long documents. 2.4 Transformer Models 2.4.1 BERT and Its Advantages Devlin et al. (2019) [ 2 ] presented the concept of the publication of BERT (Bidirectional Encoder Representations from Transformers), which was a revolution in the field of NLP potential. The major innovations of Bert as compared to its predecessors are as follows: True bidirectionality: Unlike earlier models that treated text left-to-right or concatenated separate left and right passes, BERT's masked language modeling objective enables simultaneous attention to both directions Deep contextualization: Each word's representation is a function of all other words in the input, enabling disambiguation of polysemous terms and complex financial constructions Pre-training at scale: BERT is pre-trained on 3.3 billion words, acquiring broad linguistic knowledge that transfers to specialized tasks with nominal labelled data Fine-tuning efficiency: Task-specific adaptation requires only a lightweight classification head and a few epochs of training, dramatically reducing the labeled data requirement Araci (2019) [ 1 ] demonstrated BERT's applicability to financial NLP through FinBERT, fine-tuned on the Financial PhraseBank sentiment dataset, achieving state-of-the-art accuracy of 86.2% — significantly outperforming LSTM and lexicon-based baselines. Yang et al. (2020) [ 14 ] also followed it up by further training on a 4.9 billion token financial corpus, raising the performance on various financial NLP benchmarks. Later developments In later transformer variants, such as RoBERTa (Liu et al., 2019) [ 6 ], DeBERTa (He et al., 2021) [ 3 ] and ALBERT (Lan et al., 2020) [ 5 ] added improvements to the architecture that further enhanced state-of-the-art results. The applicability of specialized pre-training on regulatory and audit text was shown by domain specific models such as SEC-BERT and AuditBERT. 2.5 Research Gap Although it has to be admitted that significant progress has been achieved, three key gaps have been identified in the current literature: Lack of context-aware models: Even when lexicon-based models or general-purpose ML models are used, it is impossible to present domain-adapted financial text in the current literature; systematic methodologies of pre-trained financial corpus and vocabulary judgement remain unrealized Limited automation: Existing scientific studies have explored individual tasks (sentiment OR risk detection OR entity recognition) singly, but do not represent a comprehensive solution of extracting insights across multiple analysis. 3. Proposed Framework / Methodology 3.1 Overview of Framework The proposed framework is an automated model of extracting business insights of unstructured financial reports. It consists of five consecutive steps, which are data collection, preprocessing, implementing the BERT model, extracting insights, and evaluating it. The pipeline will be modular too, meaning a simple stage is able to be updated or replaced, and can be scaled to massive document corpora. The input phase takes the raw financial documents and ingests them in various sources. The preprocessing phase purifies, breaks down and partakes text into BERT. The independent Bert training phase considers an adaptive task-specific flattened training of domain pre-trained approaches. Four insight module models are used in the extraction stage. The output stage presents business intelligence in the form of structured and actionable information to the consumption by the stakeholders. 3.2 Data Collection The framework ingests financial text from three primary source categories: 3.2.1 Annual Reports The major source of data is annual reports (10-K filing in the US system). These are the Management Discussion and Investigation (MD&I) section, where the rich narrative is the performance of the business, its direction and future commentary, as well as the Item 1 (Business Description), the Item 1A (Risk Factors), and the Item 7A (Quantitative and Qualitative Disclosures About Market Risk). SEC EDGAR annual reports are in machine readable XBRL format allowing automated extraction of sections. 3.2.2 Financial Disclosures Quarterly reports (10-Q) and earnings releases and call transcripts have increased frequency, which supplements annual report analysis. Transcripts of the earnings calls, including prepared statements by the management and question-and-answer portions with the analysts, are especially useful in terms of literal temperature of management reading and proactive language that is not yet contained in the official filings. 3.2.3 Public Datasets The framework is trained and evaluated using established public benchmark datasets, including the Financial PhraseBank [ 9 ] and FiQA 2018 [ 8 ] among others: Table 1 Benchmark datasets used for training and evaluation Dataset Task Size Source Financial PhraseBank Sentiment (3-class) 4,845 sent. Reuters financial news FiQA 2018 Opinion QA 17,000 sent. News & social media SEC-FLS Corpus FLS Detection 25,000 sent. SEC EDGAR 10-K filings FinNLP Risk Corpus Risk Classification 10,000 para. 10-K Item 1A sections EarningsCall QA Extractive QA 8,000 pairs S&P 500 earnings transcripts 3.3 Data Preprocessing Raw financial text requires extensive preprocessing before it can be consumed by BERT. The preprocessing pipeline operates in the following sequence: 3.3.1 Text Cleaning Financial documents contain substantial non-informative content that must be removed to improve signal quality. The cleaning stage applies the following operations: Boilerplate removal: Repeated legal disclaimers, safe harbor statements, table of contents entries, and digital signature blocks are detected and removed using regular expression patterns Table extraction: Structured numerical tables are identified and either excluded from text analysis pipelines or converted to descriptive text representations for hybrid analysis Special character normalization: Currency symbols ( $ , €, £), percentage signs, and financial shorthand ("bn", "mn", "bps") are standardized Whitespace and encoding normalization: Unicode normalization, smart quote standardization, and consistent line break handling 3.3.2 Tokenization BERT employs WordPiece subword tokenization, which decomposes words into sub-word units from a learned 30,522-entry vocabulary. This approach handles out-of-vocabulary financial terms by decomposing them into recognizable subword components. For FinBERT, the vocabulary is augmented with 500 high-frequency financial domain terms to reduce unnecessary fragmentation of important financial concepts. Special tokens are inserted according to the BERT input format: [CLS] is prepended to mark the sequence start (its final representation is used for classification tasks), and [SEP] is inserted at sentence boundaries and document segment separations. [CLS] token₁ token₂ ... tokenₙ [SEP] 3.3.3 Stopword Removal Standard English stopwords are removed from inputs destined for keyword and key phrase extraction modules. Importantly, stopword removal is not applied for sentiment analysis and classification inputs, as function words and connectors ("not", "despite", "however") carry critical sentiment-modifying information in financial text. A finance-specific stopword list supplements the standard NLTK stopword set, excluding domain-common non-informative terms such as "fiscal", "quarter", "pursuant", and "thereof". 3.4 BERT Model Implementation 3.4.1 Pre-trained BERT The backbone of the proposed framework is FinBERT, a BERT-Base model subjected to continued domain-adaptive pre-training on a 4.9 billion token financial corpus. BERT-Base comprises 12 Transformer encoder layers, 12 attention heads per layer, a hidden dimension of 768, and 110 million total parameters. The pre-training corpus for FinBERT includes Reuters financial news (1.8B tokens), SEC EDGAR 10-K/10-Q filings (2.5B tokens), and earnings call transcripts (0.6B tokens). 3.4.2 Fine-tuning Process Task-specific fine-tuning adds a lightweight classification or token-labeling head on top of the frozen-then-gradually-unfrozen BERT encoder. The fine-tuning procedure follows a three-stage protocol: Classifier warm-up (Epoch 1): Only the task-specific head is trained; BERT encoder weights are frozen. This prevents the large pre-trained representations from being disrupted by random initialization of the new head Full fine-tuning (Epochs 2–4): All layers are unfrozen and trained jointly with a low learning rate (2e-5 to 5e-5) using linear warmup over 10% of steps followed by linear decay Evaluation and selection: The checkpoint achieving best validation performance on the task-specific metric is retained Table 2 Fine-tuning hyperparameters Hyperparameter Value Optimizer AdamW Learning rate 2e-5 to 5e-5 LR schedule Linear warmup + linear decay Warmup proportion 10% of total training steps Batch size 16 / 32 Training epochs 3 to 5 Weight decay 0.01 Dropout (classifier head) 0.1 Max sequence length 128 (sentences) / 512 (paragraphs) 3.4.3 Input Representation BERT constructs its input representation by summing three distinct learned embeddings for each input token: E(t) = E_token(t) + E_segment(t) + E_position(t) Token embeddings: Map each WordPiece token to a 768-dimensional vector via a learned embedding matrix Segment embeddings: Distinguish tokens belonging to sentence A (embedding A) from sentence B (embedding B) in two-segment inputs — used for sentence-pair tasks such as next sentence prediction and question answering Positional embeddings: Encode the absolute position of each token (0 to 511) as a learned 768-dimensional vector, providing the position-agnostic Transformer with sequence order information The self-attention computation that processes these embeddings is defined as: Attention(Q, K, V) = softmax( Q Kᵀ / √d_k ) · V where Q (query), K (key), and V (value) matrices are linear projections of the input embeddings, and √d_k is a scaling factor. This mechanism allows each token to attend to all other tokens simultaneously, producing deeply contextualized representations. 3.5 Insight Extraction Modules The framework implements four specialized insight extraction modules, each implemented as a distinct fine-tuned classification or extraction head on the shared FinBERT encoder: 3.5.1 Sentiment Analysis The sentiment module will categorize the financial sentences and paragraphs into three groups, namely Positive, Neutral and Negative. A linear classification head is applied to the [CLS] token representation: P(y | x) = softmax( W_s · h_[CLS] + b_s ) where W_s ∈ ℝ^{3×768} and b_s ∈ ℝ³. The module is optimized on Financial PhraseBank and a own-generated MD&A sentiment corpus. Averaged Process sentencing level predictions are contents of sentences are aggregated to document parts with confidence[ 4 ]altered averaging data resulting in document-wide, conversation-wide, and section-wide sentiment. 3.5.2 Risk Detection Risk detection module recognizes and classifies risk-related words in financial disclosures, especially, 10-K Item 1A disclosures. It is posed as a multi-label classification scenario of seven risk categories: Market Risk: Exposure to price, interest rate, and currency fluctuations Credit Risk: Counterparty default and credit quality deterioration Liquidity Risk: Inability to meet financial obligations when due Operational Risk: Internal process failures, system outages, human error Regulatory/Legal Risk: Compliance failures, litigation, regulatory change Technology Risk: Cybersecurity, data privacy, system obsolescence Macroeconomic Risk: Recession, inflation, geopolitical instability Independent sigmoid classifiers are applied per category to the [CLS] representation, trained with binary cross-entropy loss and class-frequency weighting to handle label imbalance. 3.5.3 Key Phrase Extraction The key phrase decoding module is used to extract the most informative phrases in financial text to be summarized and indexed. It a combination of keyphrase scoring based on the importance of the tokens created by Bert with statistical keyphrase scoring: Token salience: Attention weights from the final BERT layer are aggregated to assign importance scores to individual tokens Phrase boundary detection: Adjacent high-salience tokens are merged into candidate phrases using a learned span classifier Keyphrase ranking: Candidate phrases are ranked by a combination of BERT salience score, TF-IDF weight, and position (MD&A opening paragraphs are prioritized) 3.5.4 Performance Indicators The performance indicator module derives quantitative financial values stated in narrative text, an example of which is revenue growth, EBITDA margin, EPS, return on equity, and similar sales growth. This has been applied as a Named Entity Recognition (NER) task via BIO tagging: P(y_i | x) = softmax( W_ner · h_i + b_ner ) ∀ token i Types of entities are: B-METRIC/I-METRIC metric name), B-MONEY/I-MONEY monetary value), B-PERCENT/I-PERCENT percentage figure), B-DATE/I-DATE reporting period and B-DIRECTION/I-DIRECTION increment/decrement. Obtained triplets of metric-values-period are organized in a table of performance indicators of individual documents. 3.6 Evaluation Metrics The following measures are used to determine task performance and were chosen to have a complete picture of model capability in various analytical dimensions: 3.6.1 Accuracy Measure of accuracy gives the ratio of correct classification of instances in all classes. Although intuitive, accuracy may be false when dealing with imbalanced data (e.g. when there are many Neutral sentences in sentiment classification). It is reported and accompanied with balanced measures of completeness. Accuracy = (TP + TN) / (TP + TN + FP + FN) 3.6.2 Precision The percentage of predicted positive cases, which are actually positive, is the measure of Precision - which is of crucial importance in finance as false positives (ex: a risk factor incorrectly marked as harmful) directly result into costs. Precision = TP / (TP + FP) 3.6.3 Recall Recall is the percentage of true positives correctly found - important where a false negative (including a risk factor that is missed) can be disastrous. Recall = TP / (TP + FN) 3.6.4 F1-Score The F1-score is a performance measure that gives a balanced performance of precision and recall, and it is applicable especially in the case of imbalanced financial NLP datasets. On minority classes, it is said that Macro-F1 (arithmetic mean of classes) is an indicator of their performance. F1 = 2 × (Precision × Recall) / (Precision + Recall) 4. Results and Discussion 4.1 Model Performance Comparison 4.1.1 Sentiment Analysis Results Table 3 Sentiment classification results on Financial PhraseBank dataset Model Accuracy Precision Recall F1-Score LM Dictionary Baseline 72.4% 69.3% 67.2% 68.1% Naive Bayes + TF-IDF 75.8% 72.0% 70.9% 71.4% SVM + TF-IDF 78.2% 75.6% 74.3% 74.9% LSTM + GloVe 79.6% 77.1% 76.5% 76.8% BERT-Base (fine-tuned) 88.3% 87.4% 86.5% 86.9% FinBERT — Proposed 93.6% 93.0% 91.9% 92.4% The proposed FinBERT model attains 93.6% accuracy and 92.4% F1-score which are a 21.2 percentage points higher than LM Dictionary baseline and 24.3 points higher than LM Dictionary baseline, respectively. The improvement in the best traditional baseline (SVM + TF-IDF) lies in both accuracy and F1-score at 15.4 points and 17.5 points, respectively. The comparison of BERT-Base (88.3%), with FinBERT (93.6%) shows that domain specific continued pre-training provides an extra 5.3% increase in accuracy, which proves adaptation strategy of financial corpus. 4.1.2 Risk Factor Classification Results Table 4 Risk factor classification results on FinNLP Risk Corpus 4.1.3 Named Entity Recognition (Performance Indicators) Model Accuracy Precision Recall F1-Score TF-IDF + SVM 71.2% 68.5% 66.1% 67.3% BiLSTM + Attention 78.5% 75.8% 74.2% 75.0% BERT-Base 83.7% 82.4% 81.0% 81.7% FinBERT — Proposed 91.8% 91.2% 89.7% 90.4% Table 5 Financial NER results (entity-level strict matching) Model Precision Recall F1-Score CRF + Handcrafted Features 72.1% 68.4% 70.2% BiLSTM-CRF 79.3% 76.8% 78.0% BERT-Base + NER Head 85.6% 84.2% 84.9% FinBERT + NER Head — Proposed 91.3% 90.1% 90.7% 4.2 Interpretation of Extracted Insights In addition to the aggregate measures, the quality of extracted insights is determined by qualitative model output analysis on held-out samples of financial reports. Key observations include: Sentiment nuance: FinBERT correctly classifies sentences containing financial hedging language such as "while revenues improved modestly, margin compression remains a concern" as Mixed/Negative, while baseline models assign Positive sentiment based on the word "improved" Risk categorization: The multi-label risk classifier correctly assigns Technology Risk and Operational Risk labels to cybersecurity disclosures, a distinction that single-label baselines conflate Performance indicator extraction: The NER module successfully extracts metric-value-period triplets from complex constructions such as "year-over-year revenue growth of 12.4% for the fiscal year ended December 31, 2023" with 90.7% F1 5. Implications 5.1 Managerial Implications 5.1.1 Better Decision-Making The suggested framework will convert raw financial text into structured, queryable intelligence that will be directly used to support management decision-making at various levels. Elevated executives are able to get automated corrosion of competitor MD&A trends which can foster quicker strategic reaction towards market uptakes. Without manual analyst analysis, investment committees can get risk-adjusted sentiment scores in all the portfolio company filings. In the case of board-level governance, risk factor disclosures are auto-monitored across peer companies in the industry, allowing the completeness of risk profiles to be benchmarked, and new risk themes to be identified, which are not covered by a company filing alone - a potent early warning system. Gradual decline in the tone of management is difficult to identify in a single reporting period; thus, the capability to monitor the sentiment trends of a variety of reporting periods allows for the identification of the gradual decline of the tone of the management, which could be the predecessor of earnings disappointments. 5.1.2 Risk Identification The multi-label risk detecting module is especially useful in the enterprise risk management functions. The framework makes use of automatic classification and quantification of risk disclosures among large portfolios of filers and allows: Portfolio-level risk screening: Identify all companies in a portfolio with elevated Liquidity Risk or Regulatory Risk disclosures without manual reading Risk trend analysis: Track the frequency and severity language of specific risk categories across reporting periods to detect deteriorating risk profiles Peer comparison: Benchmark a company's risk disclosure completeness against industry peers, identifying potential undisclosed risks Regulatory monitoring: Flag new regulatory risk language in real-time as 10-K filings are published on SEC EDGAR 5.2 Practical Implications 5.2.1 Automation in Financial Analytics The framework provides high productivity in financial analytics processes, which are now reliant on document-by-document review. One average financial analyst takes 3–5 hours to look through a single 10-K filing to get sentiments and risk indicators. This is brought down to minutes in the proposed framework during initial screening where the analogous analyst is able to add value to the edge cases and model output validation and more intricate synthesis that actually needs human judgment. On scale, the framework allows covering the entire universe of public company filings - more than 10,000 10-K filings filings each year in the US alone - as opposed to the small subset that manual processes can practically handle. This democratizes the access to deep financial intelligence which results in less information asymmetry between the large institutional stakeholders with the dedicated research team and smaller market participants. 5.2.2 Use in Fintech Systems The modular architecture of the proposed framework makes it well-suited for integration into fintech products and financial data platforms: Credit scoring platforms: Risk factor and sentiment signals from borrower financial disclosures can augment traditional quantitative credit models ESG analytics: Framework adaptation for ESG-related disclosure extraction supports sustainable finance analytics and regulatory reporting (e.g., TCFD climate risk disclosures) Regulatory technology (RegTech): Automated compliance monitoring of regulatory filing language supports financial institution compliance functions Investor relations platforms: Automated sentiment benchmarking of earnings communications helps IR teams calibrate messaging tone against market expectations Robo-advisory systems: Sentiment and risk signals from financial reports can inform portfolio rebalancing signals in automated investment platforms 6. Conclusion The present study suggested and tested a multifaceted BERT-based architecture to extraction of the actionable business insights out of unstructured financial reports. The framework seals an essential knowledge gap in financial analytics: the inadequacy of traditional NLP and machine learning methods to identify in-depth context-based conclusions within the intricate textual data of annual reports, MD&A sections, risk reports and earnings communications. 6.1 Summary of Findings Results of experimental outcomes of three main tasks indicate significant and consistent increases in experimental programs compared to baselines: Sentiment analysis: FinBERT achieves 93.6% accuracy on Financial PhraseBank, outperforming the best traditional baseline (SVM + TF-IDF at 78.2%) by 15.4 percentage points Risk factor classification: FinBERT achieves 91.8% accuracy on multi-label risk classification, outperforming BiLSTM-based approaches by 13.3 points Financial NER: FinBERT achieves 90.7% entity-level F1 on performance indicator extraction, outperforming BiLSTM-CRF by 12.7 points 6.2 Importance of BERT in Financial Analytics These outcomes are in agreement with the theory that the bidirectional contextual pre-training of BERT is a genuine paradigm shift of financial text analytics. It is the only model that can deal with the vagaries of financial language due to its capacity to preserve the context-dependent semantics of a word, its capacity to operate with domain-specific vocabulary and its capacity to extrapolate limited labelled data. Domain-specific continued pre-training (FinBERT) shows similar extra performance when compared to the general-domain BERT, which supports the importance of matching pre-training data distribution with domain properties. 6.3 Key Contributions The research contributes to the financial NLP literature in the following ways: An end-to-end BERT-based system, which combines sentiment analysis, risk detection, key phrase extraction, and performance indicator extraction into a single coherent system, the first such system of its kind to be described in literature A systematic comparison of B Bert and FinBert with results of lexicon-based, traditional ml, and deep learning baselines in a variety of financial NLP benchmarks, a comprehensive performance landscape Applied explanation of the business value of the framework using the case studies of financial reports, filling the gap between research in NLP and practice in the field of financial analytics A modular implementation architecture that can be extended to new financial document types, languages, and tasks of all types of analysis. 7. Future Work Although the suggested framework shows high outcomes in the target tasks, multiple significant directions are still left by the future research and development: 7.1 Real-Time Analysis The existing model is a batch model and the filings are processed after they are published. The further work will involve a creation of streaming ingestion pipeline, which follows SEC EDGAR in real-time to automatically set on it an analysis which will be performed within minutes after a new publication of a filing. This necessitates the latency of the inference time - now 45 seconds per 10 A100 GPUs - to be reduced to less than 10 seconds by model quantization (INT8), knowledge distillation (DistilFinBERT) [ 11 ], and batching. Real-time analysis would facilitate the trading signals due to time sensitivity like earnings releases and 8-K material material event disclosures. 7.2 Integration with Dashboards This structured output of the framework will be incorporated together with the gadgets of financial analytics to the interactive add-on of business intelligence interfaces to the non-technical stakeholders. Planned features include: Time-series visualization of sentiment trends across reporting periods for individual companies and industry sectors Interactive risk heat maps showing risk category intensity across portfolio companies Automated report summaries combining key phrase extraction with abstractive summarization using generative LLMs Comparative analysis tools enabling peer benchmarking of risk disclosure completeness and sentiment positioning Alert systems that notify analysts when a company's filing sentiment or risk profile deviates significantly from historical patterns or peer norms 7.3 Use of FinBERT and Advanced Models Subsequent model development will discuss a number of architectural improvements on the current implementation of FinBERT: FinBERT-Large: Scaling to the 24-layer, 340M parameter BERT-Large configuration for tasks where additional model capacity is expected to yield gains DeBERTa-Financial: Adapting DeBERTa's disentangled attention mechanism — which separately represents content and position — for financial text, where precise positional relationships between financial entities are semantically significant LLM integration: Combining the extractive precision of fine-tuned BERT models with the generative capabilities of large language models (GPT-4, Claude) for abstractive financial report summarization Cross-lingual extension: Training multilingual FinBERT variants on non-English financial corpora (EU IFRS filings, Japanese securities reports, Chinese annual reports) to support global financial analytics Temporal modeling: Incorporating longitudinal document sequences to model how company language evolves over time, enabling early detection of deteriorating financial narratives Declarations Author Contribution Dr. Prasant Kumar RoutAssistant Professor (IT), Sri Sri University, Cuttack, OdishaEmail: [email protected] : 0009-0002-4938-6170Dr. Prasant Kumar Rout contributed to the conceptualization of the study, design of the research framework, methodology development, data preprocessing, model implementation, and manuscript drafting. He played a key role in integrating the BERT-based architecture and performing the experimental analysis.Ms. Binita NandaAssistant Professor (Finance), Sri Sri University, Cuttack, OdishaEmail: [email protected] : 0000-0002-3493-4114Ms. Binita Nanda contributed to domain-specific insights, financial data interpretation, and validation of extracted business insights. She also assisted in literature review, formulation of financial context, and refinement of the results from a financial perspective.Dr. Sunil Kumar DhalProfessor (IT), Sri Sri University, Cuttack, OdishaEmail: [email protected] : 0000-0003-2324-7092Dr. Sunil Kumar Dhal provided overall supervision, guidance on research design, and critical review of the manuscript. He contributed to methodological validation, result interpretation, and final editing of the paper.All authors reviewed and approved the final version of the manuscript and agree to be accountable for all aspects of the work. References Araci, D. (2019). FinBERT: Financial sentiment analysis with pre-trained language models. arXiv preprint arXiv:1908.10063. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019, 4171–4186. He, P., Liu, X., Gao, J., & Chen, W. (2021). DeBERTa: Decoding-enhanced BERT with disentangled attention. International Conference on Learning Representations (ICLR 2021). Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., & Soricut, R. (2020). ALBERT: A lite BERT for self-supervised learning of language representations. ICLR 2020. Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692. Loughran, T., & McDonald, B. (2011). When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks. Journal of Finance, 66(1), 35–65. Maia, M., Handschuh, S., Freitas, A., Davis, B., McDermott, R., Zarrouk, M., & Balahur, A. (2018). WWW'18 open challenge: Financial opinion mining and question answering. Companion Proceedings of WWW 2018, 1941–1942. Malo, P., Sinha, A., Korhonen, P., Wallenius, J., & Takala, P. (2014). Good debt or bad debt: Detecting semantic orientations in economic texts. Journal of the Association for Information Science and Technology, 65(4), 782–796. Qin, Y., & Yang, Y. (2019). What you say and how you say it matters: Predicting stock volatility using verbal and vocal cues. Proceedings of ACL 2019, 390–401. Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108. Tetlock, P. C. (2007). Giving content to investor sentiment: The role of media in the stock market. Journal of Finance, 62(3), 1139–1168. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008. Yang, Y., Uy, M. C. S., & Huang, A. (2020). FinBERT: A pretrained language model for financial communications. arXiv preprint arXiv:2006.08097. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9209699","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":612690663,"identity":"a4b95996-05e3-47e2-b098-59d703ae5587","order_by":0,"name":"Prasant Kumar Rout","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA2UlEQVRIiWNgGAWjYHAD5gNAQkKGWOUGDAxsbAkgLTykaOExALEIazGXPvz4w88dfxjM5/d8fnWjxoKHgf3w0Q34tFj2pZlJ9p4xYJA5xrvNOucY0GE8aWk38LroDIMZA2+bAYMEG+824xw2oBYJHjMCWtg/f/wL1sLzzDjnH1FaeAykIbbwMD/ObSNCi2UPT5m0bJsxjwRbmhlzbp8EDxshv5jzsG/++LZNTk6C+fDjzznf6uT42Q8fw+8wKA2KDjYJEIsNn3JkLSDA/IGQ6lEwCkbBKBiZAAD47TvIt/I2mAAAAABJRU5ErkJggg==","orcid":"","institution":"Sri Sri University, Cuttack, Odisha","correspondingAuthor":true,"prefix":"","firstName":"Prasant","middleName":"Kumar","lastName":"Rout","suffix":""},{"id":612690664,"identity":"936219b5-09e1-4f33-aae5-41d9a27a203a","order_by":1,"name":"Binita Nanda","email":"","orcid":"","institution":"Sri Sri University, Cuttack, Odisha","correspondingAuthor":false,"prefix":"","firstName":"Binita","middleName":"","lastName":"Nanda","suffix":""},{"id":612690665,"identity":"5615420a-61ba-415e-808c-56e172cc529e","order_by":2,"name":"Sunil Kumar Dhal","email":"","orcid":"","institution":"Sri Sri University, Cuttack, Odisha","correspondingAuthor":false,"prefix":"","firstName":"Sunil","middleName":"Kumar","lastName":"Dhal","suffix":""}],"badges":[],"createdAt":"2026-03-24 09:23:51","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9209699/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9209699/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105781084,"identity":"2593c6fc-6a25-4441-a577-1d54b85bd866","added_by":"auto","created_at":"2026-03-31 05:07:48","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":84173,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eHigh-level pipeline of the proposed BERT-based financial insight extraction framework\u003c/em\u003e\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-9209699/v1/22f14063aa16aae758e571de.jpeg"},{"id":108315314,"identity":"eb21afc6-1f68-44cf-9767-948846843b75","added_by":"auto","created_at":"2026-05-02 10:25:03","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":437450,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9209699/v1/7921c045-23a0-4bf1-ad2a-f73c842502c4.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A BERT-Based Framework for Extracting Business Insights from Financial Reports","fulltext":[{"header":"1. Introduction","content":"\u003cdiv id=\"Sec2\" class=\"Section2\"\u003e \u003ch2\u003e1.1 Importance of Financial Reports in Business Decisions\u003c/h2\u003e \u003cp\u003eThe basis of corporate transparency and communication with stakeholders promote financial reporting. Annual reports, quarterly filings, management discussion and analysis (MD\u0026amp;A) applications and earnings call transcripts are all the most complete and authoritative data of the performance of a company, objectives, and second outlook. The documents will be the foremost foundation of information to a variety of parties, such as equity investors to review their positions in the portfolio, credit analysts to analyze the risks of default, regulatory authorities to ensure compliance with their regulation, and managers of the corporation to compare themselves with others in the respective industries.\u003c/p\u003e \u003cp\u003eThe strategic worth of financial reports is far more than the arranged numerical tables that they consist of. Narrative disclosures such as management commentary relating to competitive positioning, descriptions of risk factors and future guidance are likely to hold the wealthiest and most useful intelligence, even though they are no longer accessible to automated intelligence-gathering systems. Those organizations that are able to quickly and properly interpret these stories are able to have a decisive information advantage in the financial markets.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e1.2 Challenges of Analyzing Unstructured Text\u003c/h2\u003e \u003cp\u003eFinancial reports in their current form are daunting with respect to automated analysis in spite of their informational richness. An average SEC 10-K filing will be between 100 and 300 pages of formatted balance sheets and income statements which is filled with thick legal wading explanatory text. There are a number of features that make this text particularly hard to analyze:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eVolume and complexity: A single annual report may contain over 100,000 words across dozens of sections with highly specialized vocabulary\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eDomain specificity: Financial language employs technical terminology (\"amortization\", \"goodwill impairment\", \"covenant breach\") whose meaning differs substantially from everyday usage\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eContextual ambiguity: Identical phrases carry opposite sentiment in different contexts - \"above average\" can be positive (performance) or negative (risk)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eDocument heterogeneity: Reports vary significantly in structure, length, and writing style across companies, industries, and reporting periods\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eImplicit information: Critical insights are often embedded in nuanced language, hedging constructions, and comparative references that require deep comprehension to interpret\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e1.3 Rise of AI and NLP in Finance\u003c/h2\u003e \u003cp\u003eWith the intersections of progress in Natural Language Processing (NLP), and the proliferation of more digitally-published financial materials, a novel challenge of automated financial text analysis has become an opportunity of an unprecedented scale. Use of AI systems in the financial industry is being actively implemented in applications such as textual sentiment analysis of trading information, compliance filings, forecasting earnings surprises, and credit risk estimation by textual disclosure as well as news sentiment analysis to give trading signals. Transformer revolution - launched by Vaswani et al. (2017) [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] and quickly followed by the likes of BERT, GPT and RoBERTa - has dramatically increased the ability of NLP systems to reason about complex, context unique language. Naturally trained networks, which learn representations based on billions of words of text, can transfer well to the domain of particular specialized task, with relatively little fine-tuning.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e1.4 Limitations of Traditional Methods\u003c/h2\u003e \u003cp\u003eConventional techniques used to analyze financial text have been limited to rule-based and statistical techniques which are poorly adapted to the complexity of the financial language of today:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eLexicon-based methods (e.g., Loughran-McDonald dictionary): Context-blind, static, cannot adapt to new terminology or negations\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eBag-of-words models with TF-IDF: Ignore word order and syntactic structure; suffer from high dimensionality and sparsity\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSupport Vector Machines and Naive Bayes: Dependent on handcrafted features; poor generalization across document types\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eEarly neural networks (LSTM, RNN): Capture sequential structure but struggle with long-range dependencies and require large labeled datasets\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThese limitations result in systematic failures to capture the contextual, relational, and domain-specific dimensions of financial language that are most critical for insight extraction.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e1.5 Problem Statement\u003c/h2\u003e \u003cp\u003ePaper-based and other conventional automated approaches do not find rich and context-driven insights in the complex financial text. This is because of the absence of domain-specific knowledge, long document processing inability, and the inability to deal with subtle financial language leading to incomplete, inaccurate, and non-scalable financial intelligence. It is of paramount importance to have a smart, automated system that would be reliable in deriving actionable business information out of scale unstructured financial reporting\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e1.6 Objectives\u003c/h2\u003e \u003cp\u003eThe following are the main objectives to be used in this study:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eCreate a Bert-based NLP system adapted to financial report analysis\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eextraction- extract validation of major business insights on unstructured financial text, such as sentiment signals, risk factors and performance indicators\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eCompare the action of the proposed BERT-based system to the known baseline systems such as lexicon based models, conventional machine learning and previous deep-learning models\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eDiscuss practical usefulness of extracted insights to the financial stakeholders such as investors, analysts and managers\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e"},{"header":"2. Literature Review","content":" \u003cp\u003eThis section is a review of the history of research of the early financial text analytics through the traditional methods of machine learning, deep learning methods, and finally transformer- based models where the proposed framework was built.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Financial Text Analytics\u003c/h2\u003e \u003cp\u003eComputational finance of financial text has a long history that originates decades earlier. Initial studies acknowledged that predictive signatures were contained in textual information of financial reports and were not completely represented by numerical data. It is based on this foundational relationship amongst the tone of texts and stock market performance that the negative language present in Wall Street Journal columns is predictive of stock markets declines (Tetlock, 2007) [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Later endeavors were extended to the area of corporate disclosures. Loughran and McDonald (2011) [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] developed an important input by building the first lexicon of finance specific sentiment lexicon showing that negative dictionary words (using general purpose dictionary, e.g. Harvard General Inquirer) have neutral or technical senses in financial life. Their lexicon of words, such as positive, adverse, ambiguity, litigious and restricting words, was adopted as the standard of the financial NLP research almost ten years. Since then, NLP studies of finances have been developed to include various types of documents, such as quarterly earnings call transcripts (Qin and Yang, 2019) [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], analyst reports, or regulatory documents as well as social media options. These have usage in return prediction and volatility forecasting to credit risk assessment and fraud detection.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Traditional Approaches\u003c/h2\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e2.2.1 Lexicon-Based Approaches\u003c/h2\u003e \u003cp\u003eLexicon built approaches provide sentiment scores to documents by matching words in documents with predefined word lists and aggregating those scores at the sentence/document level. The most widely used lexicon that is cited in finance to gauge sentiment is Loughran- McDonald (LM) classical dictionary [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], the advanced dictionaries have since been created to process the language of the earnings calls, credit ratings statements, and central bank statements. Whereas lexicon based methods do create transparency, interpretability and are efficient in terms of computation, limitations are well documented. They are contextually blind in their essence, i.e. \"outstanding\" is a positive term whether they are speaking of an outstanding performance or outstanding liabilities. The aspects of negation (not profitable, no significant improvement), are either missing or restricted to mere negation windows. Premeditated vocabularies are not able to adjust to the changes of financial terms, changes in regulations language, and jargon usage.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e \u003ch2\u003e2.2.2 Machine Learning Approaches (SVM, Naive Bayes)\u003c/h2\u003e \u003cp\u003eSupervised machine learning models also overcame certain limitations of lexicon based methods by having classification limits learned over labelled training data and not using a set of predefined rules. TF-IDF features upon Support Vector Machines (SVMs) in the classification of financial text proved a formidable baseline in financial text classification tasks, and was proven competitive against sentiment classification and topic categorization benchmarks. Naive Bayes classifiers were based on the assumption of conditional independence and were computationally efficient and strong to small datasets, so they were used in the classification of financial news. Ensemble models such as the Random Forests and Gradient Boosting models have also been shown to achieve performance that is even better in structured sets of features that are based on both textual and non-textual cues. Nevertheless, these gains could not overcome the fact that the conventional machine learning models were limited in terms of requiring hand-crafted features, bag-of-words models that ignored the order of words in a document, and generalization of their applications across financial sectors and documents.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Deep Learning Approaches\u003c/h2\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e2.3.1 Recurrent Neural Networks (RNN)\u003c/h2\u003e \u003cp\u003eSequential Processing NLP Recurrent Neural Networks added the concept of state representations to be hidden along position of tokens in models. On financial text, RNNs also learnt longitudinal dependencies, made them become more accurately captured than bag-of-words models. Nevertheless, vanilla RNNs had problem of vanishing gradient which inhibited their capability to detect long-range dependencies essential in the interpretation of complex financial sentences.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003e2.3.2 Long Short-Term Remembrance Networks (LSTM)\u003c/h2\u003e \u003cp\u003eLong Short-Term Remembrance networks (Hochreiter \u0026amp; Schmidhuber, 1997) addressed the vanishing incline problem through fenced memory cells that selectively retain and reject information across long sequences. Bidirectional LSTMs (BiLSTMs) further extended this by processing systems in both forward and backward directions, providing richer contextual representations.\u003c/p\u003e \u003cp\u003eIn earnings call transcript prediction, Qin and Yang (2019) [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] used LSTMs to predict stocks, and reported higher performance than lexicon-based baselines. As mechanisms of attention, models that are paired with LSTMs (BiLSTM-Attention) enabled them to dynamically assign different weights to contributions of the various positions of the sentence, enabling these models to enhance their financial classification performance. Nevertheless, LSTMs still had some weaknesses such as their sequential processing limitation, ineffectiveness in being able to parallelize training and still had problems with very long documents.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Transformer Models\u003c/h2\u003e \u003cdiv id=\"Sec16\" class=\"Section3\"\u003e \u003ch2\u003e2.4.1 BERT and Its Advantages\u003c/h2\u003e \u003cp\u003eDevlin et al. (2019) [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] presented the concept of the publication of BERT (Bidirectional Encoder Representations from Transformers), which was a revolution in the field of NLP potential. The major innovations of Bert as compared to its predecessors are as follows:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eTrue bidirectionality: Unlike earlier models that treated text left-to-right or concatenated separate left and right passes, BERT's masked language modeling objective enables simultaneous attention to both directions\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eDeep contextualization: Each word's representation is a function of all other words in the input, enabling disambiguation of polysemous terms and complex financial constructions\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePre-training at scale: BERT is pre-trained on 3.3\u0026nbsp;billion words, acquiring broad linguistic knowledge that transfers to specialized tasks with nominal labelled data\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFine-tuning efficiency: Task-specific adaptation requires only a lightweight classification head and a few epochs of training, dramatically reducing the labeled data requirement\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eAraci (2019) [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] demonstrated BERT's applicability to financial NLP through FinBERT, fine-tuned on the Financial PhraseBank sentiment dataset, achieving state-of-the-art accuracy of 86.2% \u0026mdash; significantly outperforming LSTM and lexicon-based baselines. Yang et al. (2020) [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] also followed it up by further training on a 4.9\u0026nbsp;billion token financial corpus, raising the performance on various financial NLP benchmarks. Later developments In later transformer variants, such as RoBERTa (Liu et al., 2019) [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], DeBERTa (He et al., 2021) [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] and ALBERT (Lan et al., 2020) [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] added improvements to the architecture that further enhanced state-of-the-art results. The applicability of specialized pre-training on regulatory and audit text was shown by domain specific models such as SEC-BERT and AuditBERT.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Research Gap\u003c/h2\u003e \u003cp\u003eAlthough it has to be admitted that significant progress has been achieved, three key gaps have been identified in the current literature: Lack of context-aware models: Even when lexicon-based models or general-purpose ML models are used, it is impossible to present domain-adapted financial text in the current literature; systematic methodologies of pre-trained financial corpus and vocabulary judgement remain unrealized Limited automation: Existing scientific studies have explored individual tasks (sentiment OR risk detection OR entity recognition) singly, but do not represent a comprehensive solution of extracting insights across multiple analysis.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Proposed Framework / Methodology","content":"\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Overview of Framework\u003c/h2\u003e \u003cp\u003eThe proposed framework is an automated model of extracting business insights of unstructured financial reports. It consists of five consecutive steps, which are data collection, preprocessing, implementing the BERT model, extracting insights, and evaluating it. The pipeline will be modular too, meaning a simple stage is able to be updated or replaced, and can be scaled to massive document corpora.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe input phase takes the raw financial documents and ingests them in various sources. The preprocessing phase purifies, breaks down and partakes text into BERT. The independent Bert training phase considers an adaptive task-specific flattened training of domain pre-trained approaches. Four insight module models are used in the extraction stage. The output stage presents business intelligence in the form of structured and actionable information to the consumption by the stakeholders.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Data Collection\u003c/h2\u003e \u003cp\u003eThe framework ingests financial text from three primary source categories:\u003c/p\u003e \u003cdiv id=\"Sec21\" class=\"Section3\"\u003e \u003ch2\u003e3.2.1 Annual Reports\u003c/h2\u003e \u003cp\u003eThe major source of data is annual reports (10-K filing in the US system). These are the Management Discussion and Investigation (MD\u0026amp;I) section, where the rich narrative is the performance of the business, its direction and future commentary, as well as the Item 1 (Business Description), the Item 1A (Risk Factors), and the Item 7A (Quantitative and Qualitative Disclosures About Market Risk). SEC EDGAR annual reports are in machine readable XBRL format allowing automated extraction of sections.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section3\"\u003e \u003ch2\u003e3.2.2 Financial Disclosures\u003c/h2\u003e \u003cp\u003eQuarterly reports (10-Q) and earnings releases and call transcripts have increased frequency, which supplements annual report analysis. Transcripts of the earnings calls, including prepared statements by the management and question-and-answer portions with the analysts, are especially useful in terms of literal temperature of management reading and proactive language that is not yet contained in the official filings.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003e3.2.3 Public Datasets\u003c/h2\u003e \u003cp\u003eThe framework is trained and evaluated using established public benchmark datasets, including the Financial PhraseBank [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] and FiQA 2018 [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] among others:\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBenchmark datasets used for training and evaluation\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003e Dataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTask\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSize\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSource\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFinancial PhraseBank\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSentiment (3-class)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,845 sent.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eReuters financial news\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFiQA 2018\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOpinion QA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e17,000 sent.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNews \u0026amp; social media\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSEC-FLS Corpus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFLS Detection\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e25,000 sent.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSEC EDGAR 10-K filings\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFinNLP Risk Corpus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRisk Classification\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e10,000 para.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e10-K Item 1A sections\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEarningsCall QA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eExtractive QA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8,000 pairs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eS\u0026amp;P 500 earnings transcripts\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Data Preprocessing\u003c/h2\u003e \u003cp\u003eRaw financial text requires extensive preprocessing before it can be consumed by BERT. The preprocessing pipeline operates in the following sequence:\u003c/p\u003e \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e \u003ch2\u003e3.3.1 Text Cleaning\u003c/h2\u003e \u003cp\u003eFinancial documents contain substantial non-informative content that must be removed to improve signal quality. The cleaning stage applies the following operations:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eBoilerplate removal: Repeated legal disclaimers, safe harbor statements, table of contents entries, and digital signature blocks are detected and removed using regular expression patterns\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eTable extraction: Structured numerical tables are identified and either excluded from text analysis pipelines or converted to descriptive text representations for hybrid analysis\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSpecial character normalization: Currency symbols (\u003cspan\u003e$\u003c/span\u003e, \u0026euro;, \u0026pound;), percentage signs, and financial shorthand (\"bn\", \"mn\", \"bps\") are standardized\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eWhitespace and encoding normalization: Unicode normalization, smart quote standardization, and consistent line break handling\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section3\"\u003e \u003ch2\u003e3.3.2 Tokenization\u003c/h2\u003e \u003cp\u003eBERT employs WordPiece subword tokenization, which decomposes words into sub-word units from a learned 30,522-entry vocabulary. This approach handles out-of-vocabulary financial terms by decomposing them into recognizable subword components. For FinBERT, the vocabulary is augmented with 500 high-frequency financial domain terms to reduce unnecessary fragmentation of important financial concepts.\u003c/p\u003e \u003cp\u003eSpecial tokens are inserted according to the BERT input format: [CLS] is prepended to mark the sequence start (its final representation is used for classification tasks), and [SEP] is inserted at sentence boundaries and document segment separations.\u003c/p\u003e \u003cp\u003e \u003cem\u003e[CLS] token₁ token₂ ... tokenₙ [SEP]\u003c/em\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section3\"\u003e \u003ch2\u003e3.3.3 Stopword Removal\u003c/h2\u003e \u003cp\u003eStandard English stopwords are removed from inputs destined for keyword and key phrase extraction modules. Importantly, stopword removal is not applied for sentiment analysis and classification inputs, as function words and connectors (\"not\", \"despite\", \"however\") carry critical sentiment-modifying information in financial text. A finance-specific stopword list supplements the standard NLTK stopword set, excluding domain-common non-informative terms such as \"fiscal\", \"quarter\", \"pursuant\", and \"thereof\".\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003e3.4 BERT Model Implementation\u003c/h2\u003e \u003cdiv id=\"Sec29\" class=\"Section3\"\u003e \u003ch2\u003e3.4.1 Pre-trained BERT\u003c/h2\u003e \u003cp\u003eThe backbone of the proposed framework is FinBERT, a BERT-Base model subjected to continued domain-adaptive pre-training on a 4.9\u0026nbsp;billion token financial corpus. BERT-Base comprises 12 Transformer encoder layers, 12 attention heads per layer, a hidden dimension of 768, and 110\u0026nbsp;million total parameters. The pre-training corpus for FinBERT includes Reuters financial news (1.8B tokens), SEC EDGAR 10-K/10-Q filings (2.5B tokens), and earnings call transcripts (0.6B tokens).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec30\" class=\"Section3\"\u003e \u003ch2\u003e3.4.2 Fine-tuning Process\u003c/h2\u003e \u003cp\u003eTask-specific fine-tuning adds a lightweight classification or token-labeling head on top of the frozen-then-gradually-unfrozen BERT encoder. The fine-tuning procedure follows a three-stage protocol:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eClassifier warm-up (Epoch 1): Only the task-specific head is trained; BERT encoder weights are frozen. This prevents the large pre-trained representations from being disrupted by random initialization of the new head\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eFull fine-tuning (Epochs 2\u0026ndash;4): All layers are unfrozen and trained jointly with a low learning rate (2e-5 to 5e-5) using linear warmup over 10% of steps followed by linear decay\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eEvaluation and selection: The checkpoint achieving best validation performance on the task-specific metric is retained\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eFine-tuning hyperparameters\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHyperparameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eValue\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOptimizer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdamW\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLearning rate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2e-5 to 5e-5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLR schedule\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLinear warmup\u0026thinsp;+\u0026thinsp;linear decay\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWarmup proportion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10% of total training steps\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBatch size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e16 / 32\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTraining epochs\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e3 to 5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeight decay\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDropout (classifier head)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMax sequence length\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e128 (sentences) / 512 (paragraphs)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec31\" class=\"Section3\"\u003e \u003ch2\u003e3.4.3 Input Representation\u003c/h2\u003e \u003cp\u003eBERT constructs its input representation by summing three distinct learned embeddings for each input token:\u003c/p\u003e \u003cp\u003e \u003cem\u003eE(t) = E_token(t) + E_segment(t) + E_position(t)\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eToken embeddings: Map each WordPiece token to a 768-dimensional vector via a learned embedding matrix\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSegment embeddings: Distinguish tokens belonging to sentence A (embedding A) from sentence B (embedding B) in two-segment inputs \u0026mdash; used for sentence-pair tasks such as next sentence prediction and question answering\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePositional embeddings: Encode the absolute position of each token (0 to 511) as a learned 768-dimensional vector, providing the position-agnostic Transformer with sequence order information\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe self-attention computation that processes these embeddings is defined as:\u003c/p\u003e \u003cp\u003e \u003cem\u003eAttention(Q, K, V) = softmax( Q Kᵀ / \u0026radic;d_k ) \u0026middot; V\u003c/em\u003e \u003c/p\u003e \u003cp\u003ewhere Q (query), K (key), and V (value) matrices are linear projections of the input embeddings, and \u0026radic;d_k is a scaling factor. This mechanism allows each token to attend to all other tokens simultaneously, producing deeply contextualized representations.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec32\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Insight Extraction Modules\u003c/h2\u003e \u003cp\u003eThe framework implements four specialized insight extraction modules, each implemented as a distinct fine-tuned classification or extraction head on the shared FinBERT encoder:\u003c/p\u003e \u003cdiv id=\"Sec33\" class=\"Section3\"\u003e \u003ch2\u003e3.5.1 Sentiment Analysis\u003c/h2\u003e \u003cp\u003eThe sentiment module will categorize the financial sentences and paragraphs into three groups, namely Positive, Neutral and Negative. A linear classification head is applied to the [CLS] token representation:\u003c/p\u003e \u003cp\u003e \u003cem\u003eP(y | x) = softmax( W_s \u0026middot; h_[CLS] + b_s )\u003c/em\u003e \u003c/p\u003e \u003cp\u003ewhere W_s \u0026isin; ℝ^{3\u0026times;768} and b_s \u0026isin; ℝ\u0026sup3;. The module is optimized on Financial PhraseBank and a own-generated MD\u0026amp;A sentiment corpus. Averaged Process sentencing level predictions are contents of sentences are aggregated to document parts with confidence[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]altered averaging data resulting in document-wide, conversation-wide, and section-wide sentiment.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec34\" class=\"Section3\"\u003e \u003ch2\u003e3.5.2 Risk Detection\u003c/h2\u003e \u003cp\u003eRisk detection module recognizes and classifies risk-related words in financial disclosures, especially, 10-K Item 1A disclosures. It is posed as a multi-label classification scenario of seven risk categories:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eMarket Risk: Exposure to price, interest rate, and currency fluctuations\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCredit Risk: Counterparty default and credit quality deterioration\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eLiquidity Risk: Inability to meet financial obligations when due\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eOperational Risk: Internal process failures, system outages, human error\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRegulatory/Legal Risk: Compliance failures, litigation, regulatory change\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eTechnology Risk: Cybersecurity, data privacy, system obsolescence\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eMacroeconomic Risk: Recession, inflation, geopolitical instability\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eIndependent sigmoid classifiers are applied per category to the [CLS] representation, trained with binary cross-entropy loss and class-frequency weighting to handle label imbalance.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec35\" class=\"Section3\"\u003e \u003ch2\u003e3.5.3 Key Phrase Extraction\u003c/h2\u003e \u003cp\u003eThe key phrase decoding module is used to extract the most informative phrases in financial text to be summarized and indexed. It a combination of keyphrase scoring based on the importance of the tokens created by Bert with statistical keyphrase scoring:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eToken salience: Attention weights from the final BERT layer are aggregated to assign importance scores to individual tokens\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePhrase boundary detection: Adjacent high-salience tokens are merged into candidate phrases using a learned span classifier\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eKeyphrase ranking: Candidate phrases are ranked by a combination of BERT salience score, TF-IDF weight, and position (MD\u0026amp;A opening paragraphs are prioritized)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec36\" class=\"Section3\"\u003e \u003ch2\u003e3.5.4 Performance Indicators\u003c/h2\u003e \u003cp\u003eThe performance indicator module derives quantitative financial values stated in narrative text, an example of which is revenue growth, EBITDA margin, EPS, return on equity, and similar sales growth. This has been applied as a Named Entity Recognition (NER) task via BIO tagging:\u003c/p\u003e \u003cp\u003e \u003cem\u003eP(y_i | x) = softmax( W_ner \u0026middot; h_i\u0026thinsp;+\u0026thinsp;b_ner ) \u0026forall; token i\u003c/em\u003e \u003c/p\u003e \u003cp\u003eTypes of entities are: B-METRIC/I-METRIC metric name), B-MONEY/I-MONEY monetary value), B-PERCENT/I-PERCENT percentage figure), B-DATE/I-DATE reporting period and B-DIRECTION/I-DIRECTION increment/decrement. Obtained triplets of metric-values-period are organized in a table of performance indicators of individual documents.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec37\" class=\"Section2\"\u003e \u003ch2\u003e3.6 Evaluation Metrics\u003c/h2\u003e \u003cp\u003eThe following measures are used to determine task performance and were chosen to have a complete picture of model capability in various analytical dimensions:\u003c/p\u003e \u003cdiv id=\"Sec38\" class=\"Section3\"\u003e \u003ch2\u003e3.6.1 Accuracy\u003c/h2\u003e \u003cp\u003eMeasure of accuracy gives the ratio of correct classification of instances in all classes. Although intuitive, accuracy may be false when dealing with imbalanced data (e.g. when there are many Neutral sentences in sentiment classification). It is reported and accompanied with balanced measures of completeness.\u003c/p\u003e \u003cp\u003e \u003cem\u003eAccuracy = (TP\u0026thinsp;+\u0026thinsp;TN) / (TP\u0026thinsp;+\u0026thinsp;TN\u0026thinsp;+\u0026thinsp;FP\u0026thinsp;+\u0026thinsp;FN)\u003c/em\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec39\" class=\"Section3\"\u003e \u003ch2\u003e3.6.2 Precision\u003c/h2\u003e \u003cp\u003eThe percentage of predicted positive cases, which are actually positive, is the measure of Precision - which is of crucial importance in finance as false positives (ex: a risk factor incorrectly marked as harmful) directly result into costs.\u003c/p\u003e \u003cp\u003e \u003cem\u003ePrecision\u0026thinsp;=\u0026thinsp;TP / (TP\u0026thinsp;+\u0026thinsp;FP)\u003c/em\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec40\" class=\"Section3\"\u003e \u003ch2\u003e3.6.3 Recall\u003c/h2\u003e \u003cp\u003eRecall is the percentage of true positives correctly found - important where a false negative (including a risk factor that is missed) can be disastrous.\u003c/p\u003e \u003cp\u003e \u003cem\u003eRecall\u0026thinsp;=\u0026thinsp;TP / (TP\u0026thinsp;+\u0026thinsp;FN)\u003c/em\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec41\" class=\"Section3\"\u003e \u003ch2\u003e3.6.4 F1-Score\u003c/h2\u003e \u003cp\u003eThe F1-score is a performance measure that gives a balanced performance of precision and recall, and it is applicable especially in the case of imbalanced financial NLP datasets. On minority classes, it is said that Macro-F1 (arithmetic mean of classes) is an indicator of their performance.\u003c/p\u003e \u003cp\u003e \u003cem\u003eF1\u0026thinsp;=\u0026thinsp;2 \u0026times; (Precision \u0026times; Recall) / (Precision\u0026thinsp;+\u0026thinsp;Recall)\u003c/em\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4. Results and Discussion","content":"\u003cdiv id=\"Sec43\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Model Performance Comparison\u003c/h2\u003e \u003cdiv id=\"Sec44\" class=\"Section3\"\u003e \u003ch2\u003e4.1.1 Sentiment Analysis Results\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSentiment classification results on Financial PhraseBank dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLM Dictionary Baseline\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e72.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e69.3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e67.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e68.1%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNaive Bayes\u0026thinsp;+\u0026thinsp;TF-IDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e75.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e72.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e70.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e71.4%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSVM\u0026thinsp;+\u0026thinsp;TF-IDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e78.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e75.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e74.3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e74.9%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLSTM\u0026thinsp;+\u0026thinsp;GloVe\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e79.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e77.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e76.5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e76.8%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBERT-Base (fine-tuned)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e88.3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e87.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e86.5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e86.9%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFinBERT \u0026mdash; Proposed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e93.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e93.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e91.9%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e92.4%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe proposed FinBERT model attains 93.6% accuracy and 92.4% F1-score which are a 21.2 percentage points higher than LM Dictionary baseline and 24.3 points higher than LM Dictionary baseline, respectively. The improvement in the best traditional baseline (SVM\u0026thinsp;+\u0026thinsp;TF-IDF) lies in both accuracy and F1-score at 15.4 points and 17.5 points, respectively. The comparison of BERT-Base (88.3%), with FinBERT (93.6%) shows that domain specific continued pre-training provides an extra 5.3% increase in accuracy, which proves adaptation strategy of financial corpus.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec45\" class=\"Section3\"\u003e \u003ch2\u003e4.1.2 Risk Factor Classification Results\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u003cem\u003eRisk factor classification results on FinNLP Risk Corpus\u003c/em\u003e \u003cb\u003e4.1.3 Named Entity Recognition (Performance Indicators)\u003c/b\u003e\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTF-IDF\u0026thinsp;+\u0026thinsp;SVM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e71.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e68.5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e66.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e67.3%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBiLSTM\u0026thinsp;+\u0026thinsp;Attention\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e78.5%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e75.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e74.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e75.0%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBERT-Base\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e83.7%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e82.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e81.0%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e81.7%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFinBERT \u0026mdash; Proposed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e91.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e91.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e89.7%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e90.4%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eFinancial NER results (entity-level strict matching)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCRF\u0026thinsp;+\u0026thinsp;Handcrafted Features\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e72.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e68.4%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e70.2%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBiLSTM-CRF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e79.3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e76.8%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e78.0%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBERT-Base\u0026thinsp;+\u0026thinsp;NER Head\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e85.6%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e84.2%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e84.9%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFinBERT\u0026thinsp;+\u0026thinsp;NER Head \u0026mdash; Proposed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e91.3%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e90.1%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e90.7%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec46\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Interpretation of Extracted Insights\u003c/h2\u003e \u003cp\u003eIn addition to the aggregate measures, the quality of extracted insights is determined by qualitative model output analysis on held-out samples of financial reports. Key observations include:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eSentiment nuance: FinBERT correctly classifies sentences containing financial hedging language such as \"while revenues improved modestly, margin compression remains a concern\" as Mixed/Negative, while baseline models assign Positive sentiment based on the word \"improved\"\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRisk categorization: The multi-label risk classifier correctly assigns Technology Risk and Operational Risk labels to cybersecurity disclosures, a distinction that single-label baselines conflate\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePerformance indicator extraction: The NER module successfully extracts metric-value-period triplets from complex constructions such as \"year-over-year revenue growth of 12.4% for the fiscal year ended December 31, 2023\" with 90.7% F1\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"5. Implications","content":"\u003cdiv id=\"Sec48\" class=\"Section2\"\u003e \u003ch2\u003e5.1 Managerial Implications\u003c/h2\u003e \u003cdiv id=\"Sec49\" class=\"Section3\"\u003e \u003ch2\u003e5.1.1 Better Decision-Making\u003c/h2\u003e \u003cp\u003eThe suggested framework will convert raw financial text into structured, queryable intelligence that will be directly used to support management decision-making at various levels. Elevated executives are able to get automated corrosion of competitor MD\u0026amp;A trends which can foster quicker strategic reaction towards market uptakes. Without manual analyst analysis, investment committees can get risk-adjusted sentiment scores in all the portfolio company filings. In the case of board-level governance, risk factor disclosures are auto-monitored across peer companies in the industry, allowing the completeness of risk profiles to be benchmarked, and new risk themes to be identified, which are not covered by a company filing alone - a potent early warning system. Gradual decline in the tone of management is difficult to identify in a single reporting period; thus, the capability to monitor the sentiment trends of a variety of reporting periods allows for the identification of the gradual decline of the tone of the management, which could be the predecessor of earnings disappointments.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec50\" class=\"Section3\"\u003e \u003ch2\u003e5.1.2 Risk Identification\u003c/h2\u003e \u003cp\u003eThe multi-label risk detecting module is especially useful in the enterprise risk management functions. The framework makes use of automatic classification and quantification of risk disclosures among large portfolios of filers and allows:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003ePortfolio-level risk screening: Identify all companies in a portfolio with elevated Liquidity Risk or Regulatory Risk disclosures without manual reading\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRisk trend analysis: Track the frequency and severity language of specific risk categories across reporting periods to detect deteriorating risk profiles\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePeer comparison: Benchmark a company's risk disclosure completeness against industry peers, identifying potential undisclosed risks\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRegulatory monitoring: Flag new regulatory risk language in real-time as 10-K filings are published on SEC EDGAR\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec51\" class=\"Section2\"\u003e \u003ch2\u003e5.2 Practical Implications\u003c/h2\u003e \u003cdiv id=\"Sec52\" class=\"Section3\"\u003e \u003ch2\u003e5.2.1 Automation in Financial Analytics\u003c/h2\u003e \u003cp\u003eThe framework provides high productivity in financial analytics processes, which are now reliant on document-by-document review. One average financial analyst takes 3\u0026ndash;5 hours to look through a single 10-K filing to get sentiments and risk indicators. This is brought down to minutes in the proposed framework during initial screening where the analogous analyst is able to add value to the edge cases and model output validation and more intricate synthesis that actually needs human judgment. On scale, the framework allows covering the entire universe of public company filings - more than 10,000 10-K filings filings each year in the US alone - as opposed to the small subset that manual processes can practically handle. This democratizes the access to deep financial intelligence which results in less information asymmetry between the large institutional stakeholders with the dedicated research team and smaller market participants.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec53\" class=\"Section3\"\u003e \u003ch2\u003e5.2.2 Use in Fintech Systems\u003c/h2\u003e \u003cp\u003eThe modular architecture of the proposed framework makes it well-suited for integration into fintech products and financial data platforms:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eCredit scoring platforms: Risk factor and sentiment signals from borrower financial disclosures can augment traditional quantitative credit models\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eESG analytics: Framework adaptation for ESG-related disclosure extraction supports sustainable finance analytics and regulatory reporting (e.g., TCFD climate risk disclosures)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRegulatory technology (RegTech): Automated compliance monitoring of regulatory filing language supports financial institution compliance functions\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eInvestor relations platforms: Automated sentiment benchmarking of earnings communications helps IR teams calibrate messaging tone against market expectations\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRobo-advisory systems: Sentiment and risk signals from financial reports can inform portfolio rebalancing signals in automated investment platforms\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003eThe present study suggested and tested a multifaceted BERT-based architecture to extraction of the actionable business insights out of unstructured financial reports. The framework seals an essential knowledge gap in financial analytics: the inadequacy of traditional NLP and machine learning methods to identify in-depth context-based conclusions within the intricate textual data of annual reports, MD\u0026amp;A sections, risk reports and earnings communications.\u003c/p\u003e \u003cdiv id=\"Sec55\" class=\"Section2\"\u003e \u003ch2\u003e6.1 Summary of Findings\u003c/h2\u003e \u003cp\u003eResults of experimental outcomes of three main tasks indicate significant and consistent increases in experimental programs compared to baselines:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eSentiment analysis: FinBERT achieves 93.6% accuracy on Financial PhraseBank, outperforming the best traditional baseline (SVM\u0026thinsp;+\u0026thinsp;TF-IDF at 78.2%) by 15.4 percentage points\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRisk factor classification: FinBERT achieves 91.8% accuracy on multi-label risk classification, outperforming BiLSTM-based approaches by 13.3 points\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eFinancial NER: FinBERT achieves 90.7% entity-level F1 on performance indicator extraction, outperforming BiLSTM-CRF by 12.7 points\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec56\" class=\"Section2\"\u003e \u003ch2\u003e6.2 Importance of BERT in Financial Analytics\u003c/h2\u003e \u003cp\u003eThese outcomes are in agreement with the theory that the bidirectional contextual pre-training of BERT is a genuine paradigm shift of financial text analytics. It is the only model that can deal with the vagaries of financial language due to its capacity to preserve the context-dependent semantics of a word, its capacity to operate with domain-specific vocabulary and its capacity to extrapolate limited labelled data. Domain-specific continued pre-training (FinBERT) shows similar extra performance when compared to the general-domain BERT, which supports the importance of matching pre-training data distribution with domain properties.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec57\" class=\"Section2\"\u003e \u003ch2\u003e6.3 Key Contributions\u003c/h2\u003e \u003cp\u003eThe research contributes to the financial NLP literature in the following ways:\u003c/p\u003e \u003cp\u003eAn end-to-end BERT-based system, which combines sentiment analysis, risk detection, key phrase extraction, and performance indicator extraction into a single coherent system, the first such system of its kind to be described in literature\u003c/p\u003e \u003cp\u003eA systematic comparison of B Bert and FinBert with results of lexicon-based, traditional ml, and deep learning baselines in a variety of financial NLP benchmarks, a comprehensive performance landscape\u003c/p\u003e \u003cp\u003eApplied explanation of the business value of the framework using the case studies of financial reports, filling the gap between research in NLP and practice in the field of financial analytics\u003c/p\u003e \u003cp\u003eA modular implementation architecture that can be extended to new financial document types, languages, and tasks of all types of analysis.\u003c/p\u003e \u003c/div\u003e"},{"header":"7. Future Work","content":"\u003cp\u003eAlthough the suggested framework shows high outcomes in the target tasks, multiple significant directions are still left by the future research and development:\u003c/p\u003e \u003cdiv id=\"Sec59\" class=\"Section2\"\u003e \u003ch2\u003e7.1 Real-Time Analysis\u003c/h2\u003e \u003cp\u003eThe existing model is a batch model and the filings are processed after they are published. The further work will involve a creation of streaming ingestion pipeline, which follows SEC EDGAR in real-time to automatically set on it an analysis which will be performed within minutes after a new publication of a filing. This necessitates the latency of the inference time - now 45 seconds per 10 A100 GPUs - to be reduced to less than 10 seconds by model quantization (INT8), knowledge distillation (DistilFinBERT) [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], and batching. Real-time analysis would facilitate the trading signals due to time sensitivity like earnings releases and 8-K material material event disclosures.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec60\" class=\"Section2\"\u003e \u003ch2\u003e7.2 Integration with Dashboards\u003c/h2\u003e \u003cp\u003eThis structured output of the framework will be incorporated together with the gadgets of financial analytics to the interactive add-on of business intelligence interfaces to the non-technical stakeholders. Planned features include:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eTime-series visualization of sentiment trends across reporting periods for individual companies and industry sectors\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eInteractive risk heat maps showing risk category intensity across portfolio companies\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAutomated report summaries combining key phrase extraction with abstractive summarization using generative LLMs\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eComparative analysis tools enabling peer benchmarking of risk disclosure completeness and sentiment positioning\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAlert systems that notify analysts when a company's filing sentiment or risk profile deviates significantly from historical patterns or peer norms\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec61\" class=\"Section2\"\u003e \u003ch2\u003e7.3 Use of FinBERT and Advanced Models\u003c/h2\u003e \u003cp\u003eSubsequent model development will discuss a number of architectural improvements on the current implementation of FinBERT:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eFinBERT-Large: Scaling to the 24-layer, 340M parameter BERT-Large configuration for tasks where additional model capacity is expected to yield gains\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eDeBERTa-Financial: Adapting DeBERTa's disentangled attention mechanism \u0026mdash; which separately represents content and position \u0026mdash; for financial text, where precise positional relationships between financial entities are semantically significant\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eLLM integration: Combining the extractive precision of fine-tuned BERT models with the generative capabilities of large language models (GPT-4, Claude) for abstractive financial report summarization\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCross-lingual extension: Training multilingual FinBERT variants on non-English financial corpora (EU IFRS filings, Japanese securities reports, Chinese annual reports) to support global financial analytics\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eTemporal modeling: Incorporating longitudinal document sequences to model how company language evolves over time, enabling early detection of deteriorating financial narratives\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eDr. Prasant Kumar RoutAssistant Professor (IT), Sri Sri University, Cuttack, OdishaEmail: [email protected]: 0009-0002-4938-6170Dr. Prasant Kumar Rout contributed to the conceptualization of the study, design of the research framework, methodology development, data preprocessing, model implementation, and manuscript drafting. He played a key role in integrating the BERT-based architecture and performing the experimental analysis.Ms. Binita NandaAssistant Professor (Finance), Sri Sri University, Cuttack, OdishaEmail: [email protected]: 0000-0002-3493-4114Ms. Binita Nanda contributed to domain-specific insights, financial data interpretation, and validation of extracted business insights. She also assisted in literature review, formulation of financial context, and refinement of the results from a financial perspective.Dr. Sunil Kumar DhalProfessor (IT), Sri Sri University, Cuttack, OdishaEmail: [email protected]: 0000-0003-2324-7092Dr. Sunil Kumar Dhal provided overall supervision, guidance on research design, and critical review of the manuscript. He contributed to methodological validation, result interpretation, and final editing of the paper.All authors reviewed and approved the final version of the manuscript and agree to be accountable for all aspects of the work.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAraci, D. (2019). FinBERT: Financial sentiment analysis with pre-trained language models. arXiv preprint arXiv:1908.10063.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDevlin, J., Chang, M. W., Lee, K., \u0026amp; Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019, 4171\u0026ndash;4186.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe, P., Liu, X., Gao, J., \u0026amp; Chen, W. (2021). DeBERTa: Decoding-enhanced BERT with disentangled attention. International Conference on Learning Representations (ICLR 2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHochreiter, S., \u0026amp; Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735\u0026ndash;1780.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., \u0026amp; Soricut, R. (2020). ALBERT: A lite BERT for self-supervised learning of language representations. ICLR 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., \u0026amp; Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLoughran, T., \u0026amp; McDonald, B. (2011). When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks. Journal of Finance, 66(1), 35\u0026ndash;65.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaia, M., Handschuh, S., Freitas, A., Davis, B., McDermott, R., Zarrouk, M., \u0026amp; Balahur, A. (2018). WWW'18 open challenge: Financial opinion mining and question answering. Companion Proceedings of WWW 2018, 1941\u0026ndash;1942.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMalo, P., Sinha, A., Korhonen, P., Wallenius, J., \u0026amp; Takala, P. (2014). Good debt or bad debt: Detecting semantic orientations in economic texts. Journal of the Association for Information Science and Technology, 65(4), 782\u0026ndash;796.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQin, Y., \u0026amp; Yang, Y. (2019). What you say and how you say it matters: Predicting stock volatility using verbal and vocal cues. Proceedings of ACL 2019, 390\u0026ndash;401.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSanh, V., Debut, L., Chaumond, J., \u0026amp; Wolf, T. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTetlock, P. C. (2007). Giving content to investor sentiment: The role of media in the stock market. Journal of Finance, 62(3), 1139\u0026ndash;1168.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., \u0026amp; Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998\u0026ndash;6008.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, Y., Uy, M. C. S., \u0026amp; Huang, A. (2020). FinBERT: A pretrained language model for financial communications. arXiv preprint arXiv:2006.08097.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"BERT, Financial Report Analysis, NLP, Sentiment Analysis, Risk Factor Extraction, FinBERT, Business Intelligence, Transformer Models, Financial Text Analytics","lastPublishedDoi":"10.21203/rs.3.rs-9209699/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9209699/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eFinancial reports bear a lot of textual information, which is normally unstructured and imperative to business decision-making. Nevertheless, it is difficult to draw meaningful conclusions out of such huge and sophisticated papers through the conventional methods of analysis. This paper is an attempt to provide a BERT-based system to extract business insights in financial reports in an automated manner. Using the ability of Bidirectional Encoder Representations of Transformer to contextualize textual information (annual reports, management discussions, financial statements, etc.), the model proposed then analyzes the textual information to extract important information (sentiment, risk factors, performance indications, etc.).\u003c/p\u003e \u003cp\u003eThe framework includes the data preprocessing, domain-specific weighting of a pre-trained BERT, and usage of classification and information extracting methods. The research experimental results prove that BERT-based system is much more operative than the old-style machine learning (ML) models in terms of accuracy, precision and F1-score. Moreover, the inferences that were to be made are of good assistance to all stakeholders- investors, managers and analysts for making sound decisions.\u003c/p\u003e \u003cp\u003eThe paper is relevant to the area of financial text analytics research because it provides a scalable and efficient framework in processing unstructured financial information. Future development can be based on real-time analytics and connection to financial decision support systems.\u003c/p\u003e","manuscriptTitle":"A BERT-Based Framework for Extracting Business Insights from Financial Reports","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-31 05:07:45","doi":"10.21203/rs.3.rs-9209699/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e6db7cfc-5b7a-475c-b113-8a58b23076db","owner":[],"postedDate":"March 31st, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-05-02T10:24:31+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-31 05:07:45","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9209699","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9209699","identity":"rs-9209699","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00