Diachronic Profile of Startup Companies through Social Media | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Diachronic Profile of Startup Companies through Social Media Ana Rita Peixoto, Ana de Almeida, Nuno António, Fernando Batista, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2493496/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 18 Mar, 2023 Read the published version in Social Network Analysis and Mining → Version 1 posted 7 You are reading this latest preprint version Abstract Social media platforms have become powerful tools for startups, helping them find customers and raise funding. Analysing the contents posted through social media would help them make the best use of this communication and scale their business. To understand if a startup’s social media content reflects its position in its business maturation, we start by defining an adequate life cycle model for startups based on two dimensions: funding rounds and product maturity. Using Twitter as social media source of information for known Portuguese IT startups, each at their life cycle’s different phases, their tweets’ data has been analyzed. Topic modeling techniques have enabled the categorization of the data according to the topics arising in the published contents, making it possible to discover that contents can be grouped into five specific topics: “Fintech and ML”, “IT”, “Business Operations”, “Product/Service R&D”, and “Bank and Funding”. Comparing those profiles against the startup’s life cycle to understand how contents change over time provides a diachronic profile for each company. We discovered that while some topics are prevalent in the startup’s scaling, others depend on the startup’s particular phase of the cycle, revealing that startups’ Twitter social media content differs along their life cycle. Topic Modeling Social Media Startups Life Cycle Model Twitter Data Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 11 Figure 12 Figure 13 Figure 14 Figure 15 1. Introduction Social media platforms enable the creation of communities, provide easy access, and help companies promote their business. Their usage implies only a small investment driving startup companies to use it as a cost-effective tool to create a digital gateway for finding customers and raising funds. The last are two of the three critical challenges for startups reported by Wang et al. ( 2016 ), building the product is the third. These challenges derive from the fast pace of grow of a startup company, making it difficult to identify the correct steps to take for scaling-up. Gulati and DeSantola (2016) explain that startups can improve their growth and achieve their objectives by understanding the best practices of scaling. The definition of what is a startup company has evolved through time. The definition introduced by Lugović and Ahmed (2015) involves two perspectives: one concerning the business dimension and the other concerning the characteristics of the company. Regarding the business dimension, if a company has been established for less than one year and employs at least one person besides its founders, then it can be considered as a startup. As for the company characteristics, it must be an innovative and growth-oriented business. However, more recent work suggests that the startup definition depends on the actual stage of the company's life cycle (Skala 2019 ). Therefore, the startup definition is not entirely settled, but some perspectives enable the characterization of these small companies. Undoubtedly, social media has become a fundamental part of the information ecosystem, generating a large amount of data. Social media data can provide information about the clients, products, and the overall market, improving the decision-making processes. However, data has to be processed, structured, and interpreted to infer relevant decision information. Understanding social media can help improve the company’s investment (ROI) while enabling better customer relationship management (CRM), which is supported by recent studies that focus on social media data, considering it a strategic knowledge source for businesses (Kapoor et al. 2018 ). Previous studies explored digital platforms startups’ data to extract relevant information about their activity. The authors of (Saura et al. 2019), examined tweets using “#startup” to detect indicators for success and discovered the sentiment of the most common topics of tweets about startups. A broad study by Ruggieri et al. ( 2018 ) focused on finding patterns in successful startups based on their digital platforms' presence. The authors stated that newly born startups use digital platforms because it is cost-effective. Nonetheless, startups’ presence on digital platforms is continued since it enables the creation of a community between users and providers, which affects the scalability of the business, and opens new sources for creating value. Regarding the actual startups’ activity, Alotaibi et al. ( 2020 ) designed a framework to evaluate Twitter activity, using an Arabic startup as a case of study. However, to the extent of our knowledge, no study has considered the companies evolution and consequent life cycle stages. We aim to understand how a startup’s social media content changes through the different phases of its life. In other words, create a diachronic profile from the startup’s social media historical data and analyze if it reflects the startup’s scaling. Since Twitter is an ideal platform for small businesses like startups and where they are now massively present, we have chosen this platform as our primary research data source 1 . Twitter differs from other social media platforms because it gives access to a global audience where users openly communicate with other users. Above all, it offers an opportunity for businesses to interact and receive instant feedback instead of acting solely as a marketing tool (Curran et al. 2011). Thus, Twitter is a social media tool that enables a business to establish a network between customers, owners, and investors. Twitter activity is composed of tweets, which are essentially short text messages, that may include images, emoticons, URLs, mentions, and hashtags. These characteristics makes tweets categorization a challenging task. the textual analysis of startups’ tweets was performed using a topic modeling approach. We begin by assigning a category to a tweet by uncovering the tweet’s main topics and then study the evolution of tweet’s content over the startups’ life cycle. The present research differs from existent literature by linking the results of the text analysis with each company’s life cycle stage to understand if and how the startups’ social media activity alters with its rise, maturity, and consequent change of goals. This study focuses on the particular case of information technology (IT) startups founded by Portuguese executives or headquartered in Portugal as an illustrative case study. The rationale links with the fact that Portugal has created a distinctive ecosystem for IT startups over the latest years, mainly due to the Portuguese high-quality engineering talents and above-average English language fluency levels 2 . Additionally, the Portuguese government has seriously engaged in innovation policies, promoting initiatives like Startup Portugal , 200M and business incubators, which has fostered the creation of several startups. Since 2016, the investment in Lisbon-based startups has grown 30% every year 3 in a result of several successful startups and unicorns formed in Portugal. For this work, we selected eight IT startups from the Sifted 2020 Portugal startups list 4 . The chosen companies are currently at different stages in their life cycle and are considered active on Twitter. The content posted by the eight startups spans five years of analysis, from 2015 to 2020, resulting in a total of 15 577 tweets. The remainder of the paper is organized as follows: after presenting the related work, the methodology section describes our dataset, presents the methodology and proposes a new model for the life cycle of startup business. After presenting and discussing the results of our analysis in the results section, implications are drawn. To wrap up, the conclusions section describes the main conclusions and lays the path for future work. [1] https://www.thebalancesmb.com/top-reasons-why-your-small-business-should-use-twitter-2948523 [2] https://www.ef.com/wwen/epi/ [3] https://beportugal.com/startup-in-portugal/ [4] https://sifted.eu/portugal-startups-top-rankings/ 2. Related Work Social media platforms play an essential role as a marketing digital tool for small businesses. Saravanakumar and Suganthalakshmi (2012) denote social media marketing (SMM) as marketing tactics that efficiently promote brands through social media platforms. However, how can we analyze social media content and extract relevant information that demonstrates this value? This section aims to answer this question by explaining the social media analysis process and its methods and results. Additionally, we describe the startups’ life cycle since this constitutes the central hypothesis driving our research: the cycle of the startup’s life and evolution influences their social media activity. 2.1. The social media analysis process and methods Social media data has become a fundamental part of the data ecosystem and is a strategic knowledge source for decision-making (Kapoor et al. 2018 ). However, to infer relevant information from data, one needs to prepare and process it (Dutot and Mosconi 2016). Social media intelligence (SMI) collects and analyzes relevant data to provide data-driven support for strategic decisions. SMI works as a cycle because social media constantly changes, with new users creating new content and generating more data for analysis. The main focus of SMI applications is product/service reviews analysis (Kapoor et al. 2018 ). The knowledge obtained by SMI is meant to describe the present state of social media. This means that, if the objective is to predict outcomes and suggest future directions, a social media analytics (SMA) approach is deemed necessary (J. Choi et al. 2020 ). SMA and SMI present similar phases (Zeng et al. 2010 ), but the SMA methodology and results focus on the future while SMI concerns the present. Social media content is mainly text, and the goal for its analysis is to find relationships among data in textual documents and extract patterns to understand the themes being addressed (Jelodar et al. 2017 ). This can be achieved by analyzing the text’s sentiment or identifying the main topics. A topic is a list of words defined statistically to categorize the meaning of the text, and this process is termed topic modeling. In the literature, researchers address problems in the most varied fields by using topic modeling. Some of the most used methods are Latent Dirichlet Allocations (LDA) by Blei et al. ( 2002 ), whose foundations are probabilistic graphical models, Latent Semantic Analysis (LSA) by Landauer et al. ( 2007 ), and Non-negative Matrix Factorization (NMF) by Lee and Seung (2001), both based on linear algebra, namely, diverse forms of matrix factorization. LDA or Latent Dirichelet Allocation is one of the most popular and widespread methods for identification of latent topics in a text (Blei et al. 2002 ), and identifies the (relevant) topics by using generative probabilistic models. One of the areas where it is applied is in social media topic analysis, as observed in the works of (Saura et al. 2019; Yang and Zhang 2018 ; Yu et al. 2019 ). While these studies focus on different problems, each one uses topic modeling as a tool for SMA. The authors of (D. Yu et al. 2019 ) developed a novel hierarchical topic modeling technique and mined the dimension hierarchy of tweets’ topics over tweeters of different countries. Saura et al. (2019) analyzed tweets with the hashtag startup (“#startup”) and its comments. The objective was to understand which topics are present in those tweets and what are the associated sentiments. Yang and Zhang ( 2018 ) performed a similar analysis, where the authors combined topic modeling and sentiment analysis to mine the tweet’s text. They concluded that the LDA algorithm makes it easy to analyze an extensive set of tweets and obtain meaningful topics. Some other studies use topic modeling to explore and understand specific subjects on Twitter, like in the case of (Barry et al. 2018 ), which analyzes alcoholic drinks advertising, or a recent study to understand how politicians tweet about climate change by Yu et al. ( 2021 ). More recent works use topic modeling methods to examine Twitter information about COVID-19. For instance, Sha et al. ( 2020 ) analyzed governmental and politicians’ tweets about the pandemic situation and inferred a set of topics that describe Twitter activity over the countries under analysis. Kaila, R.P. & Prasad ( 2020 ) and Doogan et al. ( 2020 ) focused on tweets bearing hashtags related to COVID-19 to understand what non-government users tweet concerning the coronavirus pandemic and its global perception. While the former studies ascertain LDA as having achieved good results in analyzing Twitter posts, they also raise limitations about the use of the LDA algorithm with Twitter data. The two most common limitations are the tweets’ short text format and the need for preprocessing phase. Transforming a tweet into a document to perform a topic model might not be adequate because it has few words to extract topics. Therefore, most of the studies solve these limitations by aggregating the tweets into sets, where each collection corresponds to a document (Curiskis et al. 2020 ). However, some advances appear to avoid the aggregation, as the research by Xiong et al. ( 2018 ), where the authors proposed a short-text topic model algorithm. 2.2. Startups and social media Social media platforms have a global reach, are easy to access, and low cost, enabling startups to use social media as a digital marketing gateway and observe the market. A few studies investigate the potential relationships between startups and social media platforms in literature. Lugović and Ahmed (2015) found a positive correlation between startups’ Twitter usage and the total investment in the source country. As previously stated, Saura et al. (2019) collected tweets presenting the hashtag startup (“#startup”). The authors aimed at relating the polarity of the tweet with the topics found within the diverse sentiments. Using the tweet’s text and its comments, classified them into positive, negative, and neutral. Then, the authors performed topic modeling for each polarity and found the correspondent topics, enabling them to understand the Twitter audience sentiment of startups related content. Ruggieri et al. ( 2018 ) aimed to found patterns in successful innovative startups based on their digital platforms’ activity. Their study demonstrates that startups are present on digital platforms mostly because these platforms have a cost-effective performance. The authors also conclude that a community of users/providers of services is essential for the business. Such community is fundamental for a positive impact on digital platforms, primarily in social networking web sites, since that community provides opinions, positive or negative, about products and companies. Word-of-mouth is the everyday oral communication, which creates an impression and idea about a specific subject (Keller 2007 ) and online opinions are called electronic word-of-mouth (eWOM), as explained by Hennig-Thurau et al. ( 2004 ). Social media platforms are ideal tools for eWOM. Chu and Kim (2011) describe that eWOM enables to create a large community, which allows for the increase of digital engagement with social interactions, such as comments, likes, and shares. The last two represent non-verbal activities, and when its quantities are large, might help raising a positive feeling in the social media profile in question (Wolny and Mueller 2013). Additionally, social media activities can be used to understand the online organization's reputation (Azinhaes et al. 2021 ). 2.3. The startups’ life cycle and its stages Startups are mostly defined as fast-grow innovative businesses. According to Wang et al. ( 2016 ), the maturity evolution of a startup goes through two main stages: the learning stage and the growing stage. The learning stage consists of selecting a problem to be solved and the definition and evaluation of the solution. The problem represents a real issue or obstacle for a specific target, which is solved by providing a product or service: the solution. The product concept is developed in the growing stage, followed by an implementation start leading to a working prototype. In case it results, the startup obtains a functional product that later evolves into a mature product. However, Wang et al. ( 2016 ) emphasizes that this is not a constant cycle, saying that a startup has to go through “multiple measure-learn loops”. The loops mean a revaluation of each step as being in the stages previously referred. Concerning startups whose main product/service is software, Nguyen-Duc et al. (2015) created a conceptual model, named the hunter-gatherer, that in fact consists of two development cycles: the “hunting” cycle consists in the idea, market, and features; the “gathering” cycle features the prototype, quality, and product. The intention is that the two cycles occur at each stage, but the dimension of the cycle differs over the startup’s cycle of life. In the learning stage, the hunting cycle is more significant, while in the growing stage, the gathering cycle becomes prominent. Nevertheless, the cycles occur at each stage side-by-side: when the company obtains a mature product, the focus changes to quality matters. 3. Methodology This research follows the SMI steps framework described by Choi et al. ( 2020 ) for social media-based BI research. The process consists of four phases: “Data collection”, “Data preprocessing”, “Data analysis”, and “Validation & Interpretation”. According with this framework, the initial step was the extraction of Twitter data. As previously mentioned, a particular set of startups’ accounts was targeted: information technology (IT) startups founded by Portuguese or headquartered in Portugal, selling products or services based on machine learning (ML) approaches and presenting a B2B business model. Thus, our analysis centers in eight startups from the Sifted 2020 Portugal startups list: AttentiveMobile , Codacy , DefinedCrowd , Feedzai , Prodsmart , Talkdesk , Unbabel , and Virtuleap . After the extraction, data was cleaned and the corpus was prepared (Data preprocessing), after which we could proceed with a topic modeling (TM) technique for the analysis (Data analysis). Finally, TM results are evaluated and interpreted (Validation & Interpretation). The latter step is where the topic modeling results are compared with the startups’ funding rounds, creating a diachronic profile for each startup. For that, the funding rounds of each startup have also been collected from Crunchbase 5 , and related with the startups’ lifecycle phase. Our approach is illustrated in Fig. 1 . The features that define a startup differ depending on where in the life cycle phase the company is: in the beginning, these are innovative companies with limited resources; while in the growth process, they perform an above-average rate increment in the number of customers and revenue; and finally, they have hyper-scalability and high company valuation, which characterizes a mature startup, demonstrating that startups change over their lifecycle and that the definition of a startup depends on particular phases of company’s evolution (Skala 2019 ). Thus, startups’ life cycle is a complex concept and, as stated by Paschen ( 2017 ), it shows two different but connected perspectives that are fundamental for the company’s success: its maturity, regarding the stage of development of a product or service, and the funding rounds, that is, the fundamental investment attraction capability. Based on the related literature, we consider that the startups’ life cycle can be divided into two main perspectives. One that follows closely the concepts found in (Wang et al. 2016 ) and regards the creation of a mature product to solve a real problem: the maturity evolution . Another one concerning the startup funding rounds: the funding rounds . The funding rounds are where startups open or expose their shareholder structure to third parties, usually to business angels or venture capital firms, to secure investment to allow the startup to be able to grow (Paschen 2017 ). To illustrate a startup’s financing milestones and the startup’s evolution, we propose a life cycle model based on the previously introduced two dimensions: the funding rounds and the maturity evolution . We believe that the Funding and Product Evolution Model (FPEM) depicted in Fig. 2 , illustrates the maturation process of a startup’s life regarding time and revenue in a typical success scenario. For the model, the names of the funding rounds dimension are based on the Crunchbase Glossary 6 , and in the maturity evolution, the phases describe the startup’s product stages based on the work of Wang et al. ( 2016 ) and Paschen ( 2017 ). The proposed model, FPEM, encompasses four key phases, named after the funding round categories: the preseed phase, the seed phase, the early phase, and the late phase. For the creation of the model, we correlated the phases with the existing funding types since these are measurable, which is essential to be able to mark when a transition occurs. Then, we connected the product maturity evolution with each of the rounds. Therefore, a phase transition occurs with a funding round of a higher rank than the previous one, implying a scale-up for the company and a product maturity evolution. Typically, startups receive new funding when their product has evolved and created value for the company. However, every type of funding round can happen more than once throughout a company’s life. Notice that, for each phase, the association of concepts between maturity dimensions and funding rounds is relatively straightforward. In the preseed phase, there is only the conceptualization of a potential and innovative solution for a concrete problem. Thus, funding is usually very limited (typically below $ 150K) because it finances only an idea. These funding laps are known as angel or preseed rounds and are generally used to jump-start the company, providing financial cash to build a prototype. According to Wang et al. ( 2016 ), in this phase the startup is in its learning stage. Next, in the seed phase, a prototype or, at least, a proof-of-concept, already exists, sustaining the seed funding which can scale up to $ 2M. This round is used to build a product as market ready, incorporating the novelty proposed by the startup in the previous phase. In the early phase, the company already has a functional product and is prepared for scaling in the market. In this phase, the startup evolves for the so called growing stage (Wang et al. 2016 ). The early funding rounds, also called Series A and Series B, can have values ranging between $ 1M and $ 30M. Lastly, in the late phase, a mature product is already established and the correspondent funding, also called Series C round, usually shows values that may start at $ 10M and with no upper limit. The above-described relations between product maturity and funding rounds that represent the proposed life cycle model are validated by the topic model approach we have obtained, whose results are discussed in Section 4. The aforementioned relations enable to relate each of the four FPEM phases with the uncovered topics extracted from the tweets posted by the startups on social media during their existence. 3.1. Dataset The dataset consists of 15 577 tweets extracted from the chosen Portuguese startups’ Twitter accounts. The date of extraction date January the 10th, 2021, and the data covers every tweet posted by each of the startup since its Twitter profile creation date. The Twitter API method was employed (“GET statuses/user_timeline”) to extract all the tweets posted by providing each company account’s username, through the library tweepy (Roesslein 2020 ). The analysis focuses on the last five years, where the higher quantity of posts is concentrated: from January 2015 to December 2020, or equivalently, during 72 months. To accurately examine the startups’ activity over time, Table 1 shows the startup’s Twitter accounts’ descriptions. Table 1 Dataset description Company Name Founded Date First Tweet Date Followers Number of Tweets Tweets per month AttentiveMobile 2016 08/02/2018 1 115 695 20.44 Codacy 2012 02/10/2013 2 796 1 640 22.78 DefinedCrowd 2015 04/02/2016 1 674 1 258 21.69 Feedzai 2011 23/10/2015 2 630 3 177 51.24 Prodsmart 2012 04/12/2012 897 211 2.93 Talkdesk 2011 26/06/2019 6 586 3 211 178.39 Unbabel 2013 17/11/2013 3 510 2 615 36.32 Virtuleap 2018 29/08/2016 791 2 765 53.17 It presents the company’s first tweet available date, the number of followers, the number of tweets since January 2015, and the frequency of tweets per month. The last value regards the 72 months of analysis, or the number of months since the first tweet available date if it is more recent than January 2015. Additionally, the table shows the startup founding year, collected from Crunchbase. Figure 3 shows each startup’s quantity of tweets distributed over our chosen time window. It is possible to see that some startups post regularly, while others present peaks with more activity. Within this context, regularly means the same temporal cadence, which is the case half of the companies in analysis, namely: AttentiveMobile , DefinedCrowd , Feedzai , Talkdesk , and Unbabel. Particularly, Talkdesk account presents a higher number of tweets per month. However, not every startup presents tweet’s posts since the beginning of 2015. In the cases of AttentiveMobile , DefinedCrowd , Feedzai , and Talkdesk , the date for their first tweet available is more recent (Table 1 and Fig. 3 ). This may be due to the fact that the company’s foundation date is posterior or because more ancient tweets were voluntarily deleted. Namely, Feedzai , and Talkdesk are the ‘‘oldest’’ startups, dating from 2011, but the overall number of postings is not that high, which might suggest that they may have deleted some of their oldest tweets. Codacy , Proadsmart , and Virtuleap do not post regularly, and Virtuleap is the only company whose activity does not cover the 72 months of analysis time window. Codacy and Virtuleap present a peak in 2016 and 2017, respectively. From then on, both post with regularity but using quite fewer tweets per month. Notably, Proadsmart shows a considerable lesser degree of Twitter posting activity and is the only company who does not show posts in every month. 3.2. Text preprocessing To understand what the topics of the textual tweets might be, posted by the startups, we aggregated our dataset by month, resulting in a corpus (that is, the set of documents where each document has an id and the correspondent text) of 72 documents corresponding to each month in the time-scope of the analysis. Within each document, the id regards the month and year of the tweets. This corpus was then cleaned, retaining the vocabulary that accurately represents the startups’ content to be transformed into a document-term matrix for model training. To assure the more adequate preprocessing of tweets, we first studied the techniques applied in literature’s similar studies, thus concluding that literature supports the need for a preprocessing phase enabling as a preparation phase for achieving coherent topics. Table 2 presents the techniques that have been applied in the existent literature. Table 2 Literature preprocessing techniques usage Preprocessing technique H. J. Choi and Park (2019) Alash and Al-sultany ( 2020 ) Doogan et al. ( 2020 ) Hidayatullah et al. ( 2018 ) Yang and Zhang ( 2018 ) Lowercase transformation X X X HTML tags elimination X X X X URL elimination X X X X X Hashtag treatment X X Remove punctuation and digits X X X Remove Stop Words X X X X Lemmatization X Stemming X X N-Grams X X TF-IDF X Remove extra white spaces X X X X X Remove terms with higher frequency X X X X X Remove terms with less frequency X X X X X The most used techniques are: URL elimination, extra white spaces elimination, and exclusion of the terms presenting higher or lesser frequency, HTML tags elimination and the usage of stop words are also commonly applied. Since white spaces, URLs, and punctuation do not present information relevant towards topic’s identification, they were removed from the documents. Next, lowercase transformation and lemmatization were performed. Excluding a set of stopwords, in this case, stopwords from the Natural Language Toolkit (Bird et al. 2009 ), helps to focus the model on the relevant words that might define the text’s meaning. For this, we added the startups’ names and Twitter tags, like “RT” which means that it is a retweet, to the set of stopwords. The lemmatization goal is to convert every word to a common base form, providing coherence to the set of words and, consequently, to the topics. This was achieved via TextBlob library (Loria 2020 ). CountVectorizer from the Python library scikit-learn (Pedregosa et al. 2011 ) enables vectorizing the text and having some preprocessing customization like the use of n-grams and exclusion of terms. The n-grams used were in a range of 1 to 2, uni to bi-grams, to gather terms that may appear together, for example the bi-gram “Machine Learning”. Then, the terms that appear less than twice were excluded to prevent possible errors and misspells. Lastly, the exclusion of terms that appear in at least 80% of the tweets. Being highly frequent terms, suggests that they are meaningless in terms of topic characterization. 3.3. Topic modeling Due to its success in Twitter topic analysis related literature, the topic modeling method here employed was LDA, Latent Dirichlet Allocation (Blei et al. 2002 ). The first step is to transform the corpus into a document-term matrix, where each term is either a word or a bigram. For that, we use the frequency of the occurrence of the term/bigram in the document’s text and apply the LDA algorithm on the resulting matrix, using the Python library gensim (Rehurek et al. 2011). Since the number of topics must be given has input for the algorithm, we performed a coherence test for the advisable number of topics to be used in the modeling. Figure 4 suggests that five might be the more reliable number of topics, due to its higher coherence value. Note that the coherence measure used here was c_v , which is one of the options in gensim . Thus, the topic model created has five topics, each one characterized by the relevant terms presented in Table 3 , with all the terms showing a similar distribution within each topic. Table 3 Topic Description Topic Terms Fintech and ML future, talk, fintech, banking, reality, money2020, lisbon, project, hackathon, machinelearning Business Operations business, cloud, opentalk2020, learn, covid19, service, solution, webinar, customer service, brand Bank and Funding bank, webinar, cloud, leader, learn, read, account, report, meet, partner Product/Service RD cloud, learn, product, read, industry, innovation, boost, service, webinar, lisbon IT review, codereview, analysis, learning, websummit, machinelearning, machine learning, security, staticanalysis, lisbon The name chosen for the first topic is “Fintech and ML” because it encapsulates “fintech’’, “machine learning’’ and “banking,” as well as one event in this domain: “money2020”. The second topic is “Business Operations” since it presents terms correspondent concerns typical of the company’s operations, such as “customer service,” “brand,” “solution,” and “covid19”. Additionally, it also displays “opentalk2020”, a Talkdesk ’s event regarding customer service subjects. “Bank and Funding” is the third topic, supported by the terms “bank,” “leader,” “report,” and “partner,” while the fourth is “Product/Service R&D” sustained by terms like “innovation,” “learning,” and “boost.” Lastly, “IT” (Information Technology) is the fifth topic associated with software, like code and security, and the more significant technological event, the Websummit. [5] www.crunchbase.com [6] https://support.crunchbase.com/hc/en-us/articles/115010458467-Glossary-of-Funding-Types 4. Results And Discussion After the topic model, we divided the corpus by startup and applied the model, resulting in individual analyses representing the topics’ evolution over time for each one. In order to understand if there is a relation between the FPEM phases and the Twitter activity, we combined the funding rounds information. The first subsection describes the results obtained per startup. After the individual analysis, it became clear that there are similarities between the independent analysis, so we performed another study, using all the startups’ data, whose results are outlined in the second subsection. 4.1. Topics evolution over startups life cycle The following section regards the analysis of the Twitter activity over time for each company when combined with the startup’s funding rounds. Each figure shows the distribution of topics (in percentage), the number of tweets, and the funding rounds. To add context to the analysis, we provide, for each startup, a brief description of the company. Figure 5 represents AttentiveMobile topics’ evolution. AttentiveMobile is a B2B company that offers a personalized mobile messaging platform. We can see that from 2015 until February 2018, there are no tweets, and social media activity starts when the startup is in its early phase, where the company already has a functional product. However, the topic “Product/Service R&D” is constantly present in their tweets. In 2018, “Bank and Funding” was the topic with less presence on their content, and it increased over 2019 because they may have needed new investment. Nevertheless, “Fintech and ML” and “IT” topics are always present and are half of the content posted on Twitter, which may be related to the startup product that uses machine learning techniques. Since March 2020, the topics present a stationary distribution, with a peak of tweets quantity in June 2020. This scenario with higher activity number and stable topics’ distribution happens when the startup is in its late phase, where the company already has a mature product. The evolution of Codacy topics is depicted in Fig. 6 . Codacy is an automated code review platform. The topics distribution varies over the months, but it is clear that, in 2016, the number of tweets is significantly higher and showing two very similar peaks. During 2016, the startup is in a seed phase, which means that they should have a prototype. The startup changes to an early phase in August 2017 and, awkwardly, in September 2017, there are no tweets. The most predominant topics in its tweets are “IT”, “Fintech and ML,” and “Bank and Funding.” The first two may be related to the code review platform as it uses artificial intelligence methods, its core business, and the last may be associated with the funding needs. Next, DefinedCrowd topics’ evolution is shown in Fig. 7 . This is a company that develops artificial intelligence training data services and solutions. From 2015 until February 2016, no tweets are available, although the startup’s founding year was 2015. Until July 2018, when they received the first early round, the topics distributions fluctuate over the months, both in the number of tweets and the relative representation of topics. Once it reached its early phase, the topics started to present a more structured distribution, showing an increase of the topic “Product/Service R&D” in the tweets. According to the FPEM, this is a phase where, typically, companies own a fully functional product, justifying the increment of tweets related to “Product/Service R&D”. By the end of 2020, the graphic shows an increase in tweets per month, with two very similar peaks in July and in October. Feedzai is an artificial intelligence startup, and its core business is finance risk management. Feedzai tweet profile evolution can be observed in Fig. 8 . Notably, from January 2015 until November 2015, no tweets are available. From then on, Twitter activity starts with the company in an early phase with an already functional product. The topics show a stationary distribution, and the number of tweets is consistent over the months, except for peaks occurring in October 2017 and in October 2019, possibly because of an event occurring in October. Interestingly, in 2020 the topic “Bank and Funding” shows a decrease, and “Business Operations” has increased. The decrease may be since, in October 2017, the company reached the late phase, and raising more funds was no longer a priority. Alternatively, perhaps due to the COVID-19 ongoings, the company starts posting about the pandemic instead of financials-related tweets. Prodsmart turns factories into digital and smart ones by employing automation mechanisms into the production and control the workflow using their software. Figure 9 represents the company’s topics evolution. The tweets’ content is varied, without a visual pattern or structure, making the distribution of the topics oscillate. During 2015, April stands out with contents relating to the topic “Bank and Funding,” while in July, August, and September of the same year, the main topics in the tweets were “Product and R&D” and “Business Operations”. In terms of tweets quantity, there is a peak occurring in November 2016, although the number of tweets is always lower when compared with the other startups in the analysis. Since 2016, when the startup achieved the seed phase, the topics “Fintech and ML” and “IT”, representing the technology subject, start to be present in their tweets’ content. Over the years, the topic “Bank and Funding” has been present, which can be explained by the company’s funding needs, since throughout the time period under analysis, Prodsmart did not leave the seed phase. Figure 10 represents the Talkdesk topics’ evolution. Talkdesk is a platform to support sales teams for costumers’ satisfaction and cost savings. Although it has been founded in 2011, from 2015 until July 2019, no tweets are available. However, from July 2019 forward, the number of tweets is mostly above 100/month, which suggests that they must have deleted previous posts. From this point on, Talkdesk was at an early phase and reached the late phase in July 2020. Regarding Twitter’s activity, the topics are distributed very similarly over the months, with “Product R&D” showing the lesser number of tweets and the number of tweets showing two peaks, one in November 2019 and the other in May 2020. Since the tweets precede its entrance into a more mature phase already involving a stable product, tweeting about product development may not be between its higher priorities. Unbabel product enables companies to serve customers in their native language with a scalable translation across digital channels and Fig. 11 represents Unbabel topics’ evolution. Their first seed round was in March 2014. Therefore, Unbabel was in a seed phase until October 2016, when it reached an early phase, followed by the late phase in September 2019. In the seed phase, the tweets’ topics show an oscillatory behaviour, without a defined structure over the months. However, since the early phase, the distribution became more stationary. In September 2016, the startup shows no posts, and by September 2019, the number of tweets has decreased. Maybe because of the current late phase, they do not need to promote the product or raise more funds. Lastly, Fig. 12 depicts Virtuleap topics’ evolution. This company provides a virtual reality application that promotes brain health, with a library of games designed by neuroscientists. From 2015 to August 2016, there are no tweets available, and it is known that the company registry occurred in 2018. Virtuleap achieved a seed round in February 2018. In fact, between 2018 and 2020, the company received five seed rounds. Tweets before 2018 can be found, and there is a high value peak quantity of tweets occurring in January 2017. Additionally, the topic distribution in 2017 is stationary, being the topics “Fintech and ML” and “IT” the ones with higher representation. Since 2018, the number of posts decreases throughout the year until it reaches residual values by the last quarter of 2018. Regarding the topics, the tweet content starts to be diverse and without structure, and the ones about “Product R&D” e “Business operations” decrease. 4.2. Analysis of Twitter activity in life cycle phases The previous observations suggest that the content and the number of tweets posted by the startups may differ over their FPEM life cycle phases. In fact, it is possible to see (Fig. 13 ) that the topics percentage varies in each phase. The topic “Product R&D” is slightly higher in the preseed phase and the topic “Business Operations” is more eminent in the late phase. Newer companies focus more on product development, and mature ones already have a final product enabling them to post content about business concerns. The topics “Fintech and ML” and “IT” have similar distribution over the life cycle phases, being lower in preseed and seed phases and higher in early and late phases. Lastly, the topic “Bank and Funding” has a percentage higher than 20% in each phase, having its smaller value in the preseed and the higher in the seed. Concerning the phases, newer companies, in preseed and seed, post more about the topic “Bank and Funding”, demonstrating the importance of financials for them. In contrast, companies in the early and late phases have more content about the technology applied in their product, corresponding to the topics “Fintech and ML” and “IT”. Additionally, the preseed phase is the one with minor variance between the topics’ percentage, showing that companies in that phase may not have a specific focus for their Twitter content. For understanding if the relative emergence of topics within tweets differs according to each of the four FPEM phases, and since we have no good reason to assume that the topics distribution follows a normal distribution, the Kruskal-Wallis test was used. This is a nonparametric method that compares the means between groups, and in this scenario, each group will be one of the four life cycle phases. For the Kruskal-Wallis test, we used the SciPy (Virtanen et al. 2020 ) library, setting the significance threshold at 0.05. The null hypothesis in this scenario is that the means on each life cycle phase are the same. If the p-value is lower than the threshold, we reject the null hypothesis, meaning that the means on every life cycle are not the same. The results are presented in Table 4 , denoting the ones with a p-value below the significance with (*). Table 4 Kruskal-Wallis tests results p-value Topic: Product R&D (*) \(0.00545\) Topic: IT (*) \(2.38\text{E}-13\) Topic: Bank and Funding \(0.327\) Topic: Business Operations (*) \(2.4E-06\) Topic: Fintech and ML (*) \(8.82E-08\) Number of Tweets (*) \(2.72E-08\) (*) Statistically significant ( \(p< 0.05\) ) The topics “Product R&D”, “IT”, “Business Operations” and “Fintech and ML” present a p-value lower than the threshold, meaning that their means differ over the life cycle phases. “Bank and Funding” is the exception on the Kruskal-Wallis test, presenting a p-value expressively higher than the significance. This might implicate that it remains more stable over the life cycle, which is consistent with the analysis of the information depicted in Fig. 13 . The results also prove a statistically significant relationship between the number of tweets and the startup phases. This relationship can be visualized in Fig. 14 that shows the proportion of tweets posted per month and distributed into the life cycle phases. The graph shows that all the startups have posted more on average when traversing the preseed phase. Additionally, the higher variation in the preseed phase may be due to the fact that some of the startups in the analysis have been in this phase through a big part of the time window. However, posts from some other startups at a preseed phase were not available (or not included in the case where it occurred before 2015). Notoriously, once a seed phase is achieved, startups’ number of posts is notably less. This may be because they have received a funding round and are now more focused on product development. Nevertheless, through the early and late phases, the number of tweets slightly increases. Figure 15 displays the distributions of each topic to understand how, as shown by the Kruskal-Wallis test, do they differ throughout the FPEM phases. The topic “Product R&D” means are not equal through the life cycle phases (rejected the null hypothesis). It has higher values in the preseed phase and decreases in the subsequent ones. This change can illustrate the importance of product development in the startups’ beginning and confirms the maturity stage correspondent stated in the life cycle description of FPEM. That is, startups in the preseed phase are finding a solution to a problem. The topic “Business Operations”, which means differ over the life cycle phases, has lower values in preseed and increases over the following phases. Having the opposite behaviour of “Product R&D” and showing that with the startup growth content about product development is exchanged by business concerns. The topics “IT” and “Fintech and ML”, related to the startups’ core business in the analysis, have a similar evolution over the phases. Both topics increase until the early phase and lightly decrease in the late phase. Note that those have a statistical significance to support the means difference over the life cycle. Lastly, the topic “Bank and Funding” is the only which means do not differ over the phases, always staying around 20% value. The constant presence of this topic demonstrates the importance of fund raising and financial matters for startups and supports the fact that funding rounds are a dimension that characterizes startups. 5. Implications The primary goal in this study was to understand how Twitter contents of IT startups evolve over the company’s growth. Literature shows that startups experience characteristic phases due to companies changing through their life cycle, adjusting their goals. The first contribution of this study is the conceptualization of a life cycle model. This proposal is based on two dimensions previously described in the literature: maturity evolution and funding rounds. Maturity regards the development of the product or service that the startup is selling, and the funding rounds regards capitalization through investors financing. Our proposal unites those dimensions, creating a natural flow of business evolution: the Funding and Product Evolution Model (FPEM). The second important implication of this study is the categorization of IT startups’ social media activity. Understanding the Twitter content was achieved through topic modeling, leading to a well defined set of five topics describing the main subjects in the startups’ tweets, which are: “Fintech and ML”; “IT”; “Business Operations”; “Product/Service R&D”; and “Bank and Funding”. The third implication brings light into the question of how the startup’s phases within its life cycle may affect social media usage. Our findings suggest that Twitter content produced by IT startups changes over the FPEM phases while the startups scale-up. The results outline that startups’ initial posts are primarily related to product development and, in more advanced maturity phases, tweets became related to operations and business concerns. As expected, one of the topics found, “Bank and Funding”, constantly emerges in tweets over the entire life cycle, denoting financial matters are a cornerstone for startups, as it should be expected due to the particularities of these companies. 6. Conclusions And Future Work This study proposes a new startup’s life cycle model based on funding rounds and the companies product maturity: the Funding and Product Evolution Model - FPEM. The validity of FPEM is illustrated using an SMI cycle-based methodology to extract the main topics from eight IT startups founded by Portugueses or headquartered in Portugal. The Twitter posts were subjected to an automatic information extraction of topics to understand if the tweets’ contents change while startups are scaling up. For the IT startups chosen, the tweets posted between 2015 and 2020 were subjected to a topic model analysis, adding up to a total of selected 15 577 tweets. The results were combined with the FPEM life cycle model, creating a diachronic profile for each one of the startups. It was possible to perceive that the startups’ key topics are: “Fintech and ML” and “IT” which regard the startups’ core business; “Business Operations” and “Product/Service R&D” about enterprise subjects and product development; “Bank and Funding” concerning startups’ financing. Nevertheless, results reveal that IT startups’ Twitter topics change over time according to the company’s current cycle of life. The number of tweets published also varies according to the startup phase, showing that newer and more mature IT startups post more on Twitter when compared to companies in an intermediate phase. In terms of content, the topic “Bank and Funding” is the only one of the five topics present throughout startups’ life cycle, demonstrating the great importance of financial investments and capital enabling the company’s growth. On the other hand, another uncovered topic, “Product R&D”, is predominant during the preseed phase, showing that startups begin as product-focused companies. In contrast, the topic “Business Operations” is prevalent in the late phase, revealing that business concerns take the place of the product development content with the startup’s growth. Therefore, social media content evolves with the startups’ evolution and consequently with the scaling stages. Like all studies, this one is not without limitations. The first is the fact that it focused only on Portugal-based (or related) IT startups. Future research should study startups from other industries and countries to confirm if the results are similar, independently of the industry and region. This study relies only on publicly available Twitter data. Future studies should use data from other social media platforms, like LinkedIn, to understand if posted contents vary for different platforms or if complementary topics emerge. In this study, the startups were at different phases that limited the possibility of a complete startup life cycle for some. Future research could focus on studying other startups at the same phase of the FPEM for more comprehensive results. Lastly, we only validated the FPEM with the topics extracted from social media. Future work must use other data sources concerning startups to revalidate the model, like interviews with startups’ founders and venture capital experts. Declarations We have no known conflict of interest to disclose. This work was partly funded through national funds by FCT - Fundação para a Ciência e Tecnologia, I.P. under projects UIDB/04466/2020 (ISTAR), and UIDB/50021/2020 (INESC-ID). Author Notes Ana Rita Peixoto https://orcid.org/0000-0001-7618-5994 Ana de Almeida https://orcid.org/0000-0001-9519-4634 Nuno António https://orcid.org/0000-0002-4801-2487 Fernando Batista https://orcid.org/0000-0002-1075-0177 Ricardo Ribeiro https://orcid.org/0000-0002-2058-693X References Alash, Hayder M, and Ghaidaa A Al-sultany. 2020. “Improve Topic Modeling Algorithms Based on Twitter Hashtags Improve Topic Modeling Algorithms Based on Twitter Hashtags.” Alotaibi, Bashayer et al. 2020. “Startup Initiative Response Analysis (SIRA) Framework for Analyzing Startup Initiatives on Twitter.” IEEE Access 8: 10718–30. Azinhaes, J., F. Batista, and J. C. Ferreira. 2021. “EWOM for Public Institutions: Application to the Case of the Portuguese Army.” Social Network Analysis and Mining 11(1). https://doi.org/10.1007/s13278-021-00837-w. Barry, Adam E., Danny Valdez, Alisa A. Padon, and Alex M. Russell. 2018. “Alcohol Advertising on Twitter—A Topic Model.” American Journal of Health Education 49(4): 256–63. https://doi.org/10.1080/19325037.2018.1473180. Bird, Steven., Ewan. Klein, and Edward. Loper. 2009. Natural language processing with Python Natural Language Processing with Python . O’Reilly. https://www.oreilly.com/library/view/natural-language-processing/9780596803346/ (January 16, 2023). Blei, David M., Andrew Y. Ng, and Michael T. Jordan. 2002. “Latent Dirichlet Allocation.” Advances in Neural Information Processing Systems 3: 993–1022. Choi, Hyeok Jun, and Cheong Hee Park. 2019. “Emerging Topic Detection in Twitter Stream Based on High Utility Pattern Mining.” Expert Systems with Applications 115: 27–36. https://doi.org/10.1016/j.eswa.2018.07.051. Choi, Jaewoong et al. 2020. “Social Media Analytics and Business Intelligence Research: A Systematic Review.” Information Processing and Management 57(6): 102279. https://doi.org/10.1016/j.ipm.2020.102279. Chu, Shu-Chuan, and Yoojung Kim. 2011. “Determinants of Consumer Engagement in Electronic Word-of-Mouth (EWOM) in Social Networking Sites.” International Journal of Advertising 30(1): 47–75. https://www.tandfonline.com/doi/full/10.2501/IJA-30-1-047-075. Curiskis, Stephan A., Barry Drake, Thomas R. Osborn, and Paul J. Kennedy. 2020. “An Evaluation of Document Clustering and Topic Modelling in Two Online Social Networks: Twitter and Reddit.” Information Processing and Management 57(2): 102034. https://doi.org/10.1016/j.ipm.2019.04.002. Curran, Kevin, Kevin O’Hara, and Sean O’Brien. 2011. “The Role of Twitter in the World of Business.” International Journal of Business Data Communications and Networking 7(3): 1–15. Doogan, Caitlin, Wray Buntine, Henry Linger, and Samantha Brunt. 2020. “Public Perceptions and Attitudes Toward COVID-19 Nonpharmaceutical Interventions Across Six Countries: A Topic Modeling Analysis of Twitter Data.” Journal of medical Internet research 22(9): e21419. Dutot, Vincent, and Elaine Mosconi. 2016. “Social Media and Business Intelligence: Defining and Understanding Social Media Intelligence.” Journal of Decision Systems 25(3): 191–92. Gulati, Ranjay, and Alicia DeSantola. 2016. “Start-Ups That Last.” Harvard Business Review 2016(March). https://hbr.org/2016/03/start-ups-that-last (January 16, 2023). Hennig-Thurau, Thorsten, Kevin P. Gwinner, Gianfranco Walsh, and Dwayne D. Gremler. 2004. “Electronic Word-of-Mouth via Consumer-Opinion Platforms: What Motivates Consumers to Articulate Themselves on the Internet?” Journal of Interactive Marketing 18(1): 38–52. https://linkinghub.elsevier.com/retrieve/pii/S1094996804700961. Hidayatullah, Ahmad Fathan et al. 2018. “Twitter Topic Modeling on Football News.” 2018 3rd International Conference on Computer and Communication Systems, ICCCS 2018 : 94–98. Jelodar, Hamed et al. 2017. “Latent Dirichlet Allocation (LDA) and Topic Modeling: Models, Applications, a Survey.” Multimedia Tools and Applications 78: 183–98. http://arxiv.org/abs/1711.04305. Kaila, R.P. & Prasad, A.V.K. 2020. “Informational Flow on Twitter - Corona Virus Outbreak – Topic.” 11(3): 128–34. Kapoor, Kawaljeet Kaur et al. 2018. “Advances in Social Media Research: Past, Present and Future.” Information Systems Frontiers 20(3): 531–58. Keller, Ed. 2007. “Unleashing the Power of Word of Mouth: Creating Brand Advocacy to Drive Growth.” Journal of Advertising Research 47(4): 448–52. Landauer, Thomas K and McNamara, Danielle S and Dennis, Simon and Kintsch, Walter. 2007. Handbook of latent semantic analysis. Handbook of Latent Semantic Analysis. Psychology Press. Lee, Daniel D, and H Sebastian Seung. 2001. “Algorithms for Non-Negative Matrix Factorization.” Advances in Neural Information Processing Systems 13. https://proceedings.neurips.cc/paper/2000/file/f9d1152547c0bde01830b7e8bd60024c-Paper.pdf (January 16, 2023). Loria, Steven. 2020. “TextBlob: Simplified Text Processing — TextBlob 0.16.0 Documentation.” https://textblob.readthedocs.io/en/dev/ (January 16, 2023). Lugović, Sergej, and Wasim Ahmed. 2015. “An Analysis of Twitter Usage Among Startups in Europe.” In , 299–308. http://infoz.ffzg.hr/infuture/2015/images/papers/8-02 Lugovic, Ahmed, An Analysis of Twitter Usage Among Startups in EU.pdf. Nguyen-Duc, Anh, Pertti Seppänen, and Pekka Abrahamsson. 2015. “Hunter-Gatherer Cycle: A Conceptual Model of the Evolution of Software Startups.” ACM International Conference Proceeding Series 24-26-Augu(Idi): 199–203. Paschen, Jeannette. 2017. “Choose Wisely: Crowdfunding through the Stages of the Startup Life Cycle.” Business Horizons 60(2): 179–88. http://dx.doi.org/10.1016/j.bushor.2016.11.003. Pedregosa, Fabian et al. 2011. 12 Journal of Machine Learning Research Scikit-Learn: Machine Learning in Python . http://scikit-learn.sourceforge.net. (January 16, 2023). Rehurek, Radim and Sojka, Petr. 2011. “Gensim--Python Framework for Vector Space Modelling.” NLP Centre, Faculty of Informatics, Masaryk University, Brno, Czech Republic 3. Roesslein, Joshua. 2020. “Tweepy: Twitter for Python!” https://github.com/tweepy/tweepy (January 16, 2023). Ruggieri, Roberto et al. 2018. “The Impact of Digital Platforms on Business Models: An Empirical Investigation on Innovative Start-Ups.” Management and Marketing 13(4): 1210–25. Saravanakumar, M, and T Suganthalakshmi. 2012. “Social Media Marketing.” Life Science Journal 9(4): 1097–8135. http://www.lifesciencesite.comhttp//www.lifesciencesite.com.670 (January 16, 2023). Saura, Jose Ramon, Pedro Palos-Sanchez, and Antonio Grilo. 2019. “Detecting Indicators for Startup Business Success: Sentiment Analysis Using Text Data Mining.” Sustainability (Switzerland) 11(3): 1–14. Sha, Hao, Mohammad Al Hasan, George Mohler, and P. Jeffrey Brantingham. 2020. “Dynamic Topic Modeling of the COVID-19 Twitter Narrative among U.S. Governors and Cabinet Executives.” arXiv (2): 2–7. http://arxiv.org/abs/2004.11692. Skala, Agnieszka. 2019. Digital Startups in Transition Economies Digital Startups in Transition Economies . Virtanen, Pauli et al. 2020. “SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python.” Nature Methods 17(3): 261–72. http://www.nature.com/articles/s41592-019-0686-2. Wang, Xiaofeng et al. 2016. “Key Challenges in Software Startups across Life Cycle Stages.” Lecture Notes in Business Information Processing 251: 169–82. Wolny, Julia, and Claudia Mueller. 2013. “Analysis of Fashion Consumers’ Motives to Engage in Electronic Word-of-Mouth Communication through Social Media Platforms.” Journal of Marketing Management 29(5–6): 562–83. Xiong, Shufeng, Kuiyi Wang, Donghong Ji, and Bingkun Wang. 2018. “A Short Text Sentiment-Topic Model for Product Reviews.” Neurocomputing 297: 94–102. https://doi.org/10.1016/j.neucom.2018.02.034. Yang, Sidi, and Haiyi Zhang. 2018. “Text Mining of Twitter Data Using a Latent Dirichlet Allocation Topic Model and Sentiment Analysis.” International Journal of Computer and Information Engineering 12(7): 525–29. Yu, Chao et al. 2021. “Tweeting About Climate: Which Politicians Speak Up and What Do They Speak Up About?” Social Media + Society 7(3): 205630512110338. http://journals.sagepub.com/doi/10.1177/20563051211033815. Yu, Dongjin, Dengwei Xu, Dongjing Wang, and Zhiyong Ni. 2019. “Hierarchical Topic Modeling of Twitter Data for Online Analytical Processing.” IEEE Access 7: 12373–85. Zeng, Daniel, Hsinchun Chen, Robert Lusch, and Shu Hsing Li. 2010. “Social Media Analytics and Intelligence.” IEEE Intelligent Systems 25(6): 13–16. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 18 Mar, 2023 Read the published version in Social Network Analysis and Mining → Version 1 posted Editorial decision: Major revision 05 Feb, 2023 Reviews received at journal 30 Jan, 2023 Reviewers agreed at journal 23 Jan, 2023 Reviewers invited by journal 23 Jan, 2023 Editor assigned by journal 23 Jan, 2023 Submission checks completed at journal 19 Jan, 2023 First submitted to journal 18 Jan, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2493496","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":168993580,"identity":"b814975e-6360-4c3f-b1b5-af408dc5775f","order_by":0,"name":"Ana Rita Peixoto","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABA0lEQVRIie2PsWrDMBCGTxzYy4FXlQT8ChcCbQYTv4pLoFNIA1k6mXRJFkPWDH2QjDIasvgBOiZkzeAxQ0sr1aST5ayF6hukE+K7nx/A4/mLYHPRz1lz81Iw71LwqiCILTeyAr5u6YixA9JvXofysI7KWuySPu835SmZ53EaLoUmhlnqUPoaUYrqie4KjcMp60FBCqyycKVI00WKlaZIToLelFVG8rm2ymPhVvAiVl8UWGXEeUbxAW4pgUlRTQowmhS4qdyPRDWxXYaDwnapMijfWC5Mp3YlKk/vYjdOef96PFw+8jhcK6zPL8ksXDpiDPjZ8ikzt2AQbfu6FY/H4/lPfAPniEZTrzbG2AAAAABJRU5ErkJggg==","orcid":"","institution":"Instituto Universitário de Lisboa (ISCTE-IUL)","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Ana","middleName":"Rita","lastName":"Peixoto","suffix":""},{"id":168993581,"identity":"bec76ca3-69d4-46d1-8105-7a4bb146e855","order_by":1,"name":"Ana de Almeida","email":"","orcid":"","institution":"Instituto Universitário de Lisboa (ISCTE-IUL)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ana","middleName":"","lastName":"de Almeida","suffix":""},{"id":168993582,"identity":"e9fcd7f2-c3c2-4519-b0c4-7538ab51d5a4","order_by":2,"name":"Nuno António","email":"","orcid":"","institution":"Nova Information Management School (NOVA IMS), Universidade Nova de Lisboa","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Nuno","middleName":"","lastName":"António","suffix":""},{"id":168993583,"identity":"d20c5f68-ae96-4fb4-9513-2d8ccd20946d","order_by":3,"name":"Fernando Batista","email":"","orcid":"","institution":"Instituto Universitário de Lisboa (ISCTE-IUL)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Fernando","middleName":"","lastName":"Batista","suffix":""},{"id":168993584,"identity":"7b7d087b-80bf-49c8-b4a9-efd614e1eb2e","order_by":4,"name":"Ricardo Ribeiro","email":"","orcid":"","institution":"Instituto Universitário de Lisboa (ISCTE-IUL)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ricardo","middleName":"","lastName":"Ribeiro","suffix":""}],"badges":[],"createdAt":"2023-01-19 00:44:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2493496/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2493496/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s13278-023-01055-2","type":"published","date":"2023-03-18T20:02:47+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":31869167,"identity":"af2f6f69-8e2b-4d5f-9fd0-875b7c6c8e90","added_by":"auto","created_at":"2023-01-20 16:04:46","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":70321,"visible":true,"origin":"","legend":"\u003cp\u003eProject Pipeline\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/43218f54a665f4bb5c86cfb5.png"},{"id":31869168,"identity":"a892d2b0-b2ba-4709-93b0-240dfb5d2a4b","added_by":"auto","created_at":"2023-01-20 16:04:47","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":73634,"visible":true,"origin":"","legend":"\u003cp\u003eStartups’ life cycle model - Funding and Product Evolution Model (FPEM)\u003c/p\u003e","description":"","filename":"Figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/e97157449c1cfbf4fb341b9a.jpg"},{"id":31869796,"identity":"160eed1e-8971-463f-bab5-951164f93547","added_by":"auto","created_at":"2023-01-20 16:12:47","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":102360,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of tweets quantity over time\u003c/p\u003e","description":"","filename":"Figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/e0e170401cfcf0e7298d32d8.jpg"},{"id":31870844,"identity":"847ba6a6-14f2-4438-bd6b-dd69d5650292","added_by":"auto","created_at":"2023-01-20 16:36:47","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":24354,"visible":true,"origin":"","legend":"\u003cp\u003eLDA coherence analysis\u003c/p\u003e","description":"","filename":"Figure4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/82c60786e0bf5397e90ae0c5.jpg"},{"id":31870347,"identity":"660bef44-4650-4871-bff8-73c5ed3f6edd","added_by":"auto","created_at":"2023-01-20 16:20:47","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":233873,"visible":true,"origin":"","legend":"\u003cp\u003eAttentive Mobile\u003c/p\u003e","description":"","filename":"Figure5attentive.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/d4c89cc823a6e7852e1e888f.jpeg"},{"id":31870505,"identity":"df5663ca-45a8-462a-b730-bdfd69da19ad","added_by":"auto","created_at":"2023-01-20 16:28:47","extension":"jpeg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":312557,"visible":true,"origin":"","legend":"\u003cp\u003eCodacy\u003c/p\u003e","description":"","filename":"Figure6codacy.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/bb71d0ce02162b9d4be38a12.jpeg"},{"id":31869804,"identity":"146f597c-1ffb-4bee-8197-277f5985b370","added_by":"auto","created_at":"2023-01-20 16:12:47","extension":"jpeg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":296834,"visible":true,"origin":"","legend":"\u003cp\u003eDefined Crowd\u003c/p\u003e","description":"","filename":"Figure7definedcrowd.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/ea417df2098dacb505665b32.jpeg"},{"id":31870354,"identity":"77169655-814a-4fac-8b6e-31fcd60f9d68","added_by":"auto","created_at":"2023-01-20 16:20:47","extension":"jpeg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":289355,"visible":true,"origin":"","legend":"\u003cp\u003eFeedzai\u003c/p\u003e","description":"","filename":"Figure8feedzai.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/add96dedbb92fdd98d5a8ba2.jpeg"},{"id":31870510,"identity":"7f5ecaad-afc5-4f87-bad2-623f1624abb9","added_by":"auto","created_at":"2023-01-20 16:28:47","extension":"jpeg","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":316258,"visible":true,"origin":"","legend":"\u003cp\u003eUnbabel\u003c/p\u003e","description":"","filename":"Figure11unbabel.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/a9f92879f3f3ab20f6a06123.jpeg"},{"id":31869797,"identity":"70138391-b69f-44ee-a7dd-261a511b6c7f","added_by":"auto","created_at":"2023-01-20 16:12:47","extension":"jpeg","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":271335,"visible":true,"origin":"","legend":"\u003cp\u003eVirtuleap\u003c/p\u003e","description":"","filename":"Figure12virtuleap.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/14be3ef8634918faf4b9a478.jpeg"},{"id":31870845,"identity":"26101b87-ceba-4fa4-beec-9b541f296ac0","added_by":"auto","created_at":"2023-01-20 16:36:47","extension":"jpg","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":67122,"visible":true,"origin":"","legend":"\u003cp\u003eAverage of the topics’ predominance per phase\u003c/p\u003e","description":"","filename":"Figure13average.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/c08a26147e595c17eff60743.jpg"},{"id":31870848,"identity":"8a6335e6-e8e3-42dd-a54a-5efec728e46a","added_by":"auto","created_at":"2023-01-20 16:36:47","extension":"jpg","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":18677,"visible":true,"origin":"","legend":"\u003cp\u003eTweets quantity over life cycle phases\u003c/p\u003e","description":"","filename":"Figure14.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/097ec7f7a60281def3b53fdf.jpg"},{"id":31869181,"identity":"097110b7-d444-487f-adb5-427980c6f34d","added_by":"auto","created_at":"2023-01-20 16:04:47","extension":"jpg","order_by":15,"title":"Figure 15","display":"","copyAsset":false,"role":"figure","size":142271,"visible":true,"origin":"","legend":"\u003cp\u003eTopics distribution over life cycle phases\u003c/p\u003e","description":"","filename":"Figure15.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/f1787651daaef8d65c43e1b1.jpg"},{"id":44722904,"identity":"7a1c61b7-d2b2-4c4c-9caf-2bb431df9e23","added_by":"auto","created_at":"2023-10-16 20:09:10","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1368277,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2493496/v1/2916fa23-d08c-43f7-95bb-e837884a7eca.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Diachronic Profile of Startup Companies through Social Media","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eSocial media platforms enable the creation of communities, provide easy access, and help companies promote their business. Their usage implies only a small investment driving startup companies to use it as a cost-effective tool to create a digital gateway for finding customers and raising funds. The last are two of the three critical challenges for startups reported by Wang et al. (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2016\u003c/span\u003e), building the product is the third. These challenges derive from the fast pace of grow of a startup company, making it difficult to identify the correct steps to take for scaling-up. Gulati and DeSantola (2016) explain that startups can improve their growth and achieve their objectives by understanding the best practices of scaling.\u003c/p\u003e \u003cp\u003eThe definition of what is a startup company has evolved through time. The definition introduced by Lugović and Ahmed (2015) involves two perspectives: one concerning the business dimension and the other concerning the characteristics of the company. Regarding the business dimension, if a company has been established for less than one year and employs at least one person besides its founders, then it can be considered as a startup. As for the company characteristics, it must be an innovative and growth-oriented business. However, more recent work suggests that the startup definition depends on the actual stage of the company's life cycle (Skala \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Therefore, the startup definition is not entirely settled, but some perspectives enable the characterization of these small companies.\u003c/p\u003e \u003cp\u003eUndoubtedly, social media has become a fundamental part of the information ecosystem, generating a large amount of data. Social media data can provide information about the clients, products, and the overall market, improving the decision-making processes. However, data has to be processed, structured, and interpreted to infer relevant decision information. Understanding social media can help improve the company\u0026rsquo;s investment (ROI) while enabling better customer relationship management (CRM), which is supported by recent studies that focus on social media data, considering it a strategic knowledge source for businesses (Kapoor et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). Previous studies explored digital platforms startups\u0026rsquo; data to extract relevant information about their activity. The authors of (Saura et al. 2019), examined tweets using \u0026ldquo;#startup\u0026rdquo; to detect indicators for success and discovered the sentiment of the most common topics of tweets about startups. A broad study by Ruggieri et al. (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) focused on finding patterns in successful startups based on their digital platforms' presence. The authors stated that newly born startups use digital platforms because it is cost-effective. Nonetheless, startups\u0026rsquo; presence on digital platforms is continued since it enables the creation of a community between users and providers, which affects the scalability of the business, and opens new sources for creating value. Regarding the actual startups\u0026rsquo; activity, Alotaibi et al. (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) designed a framework to evaluate Twitter activity, using an Arabic startup as a case of study. However, to the extent of our knowledge, no study has considered the companies evolution and consequent life cycle stages. We aim to understand how a startup\u0026rsquo;s social media content changes through the different phases of its life. In other words, create a diachronic profile from the startup\u0026rsquo;s social media historical data and analyze if it reflects the startup\u0026rsquo;s scaling.\u003c/p\u003e \u003cp\u003eSince Twitter is an ideal platform for small businesses like startups and where they are now massively present, we have chosen this platform as our primary research data source\u003ca class=\"FNLink\" href=\"#Fn1\" id=\"#FNLinkFn1\"\u003e1\u003c/a\u003e. Twitter differs from other social media platforms because it gives access to a global audience where users openly communicate with other users. Above all, it offers an opportunity for businesses to interact and receive instant feedback instead of acting solely as a marketing tool (Curran et al. 2011). Thus, Twitter is a social media tool that enables a business to establish a network between customers, owners, and investors. Twitter activity is composed of tweets, which are essentially short text messages, that may include images, emoticons, URLs, mentions, and hashtags. These characteristics makes tweets categorization a challenging task. the textual analysis of startups\u0026rsquo; tweets was performed using a topic modeling approach. We begin by assigning a category to a tweet by uncovering the tweet\u0026rsquo;s main topics and then study the evolution of tweet\u0026rsquo;s content over the startups\u0026rsquo; life cycle. The present research differs from existent literature by linking the results of the text analysis with each company\u0026rsquo;s life cycle stage to understand if and how the startups\u0026rsquo; social media activity alters with its rise, maturity, and consequent change of goals.\u003c/p\u003e \u003cp\u003eThis study focuses on the particular case of information technology (IT) startups founded by Portuguese executives or headquartered in Portugal as an illustrative case study. The rationale links with the fact that Portugal has created a distinctive ecosystem for IT startups over the latest years, mainly due to the Portuguese high-quality engineering talents and above-average English language fluency levels\u003ca class=\"FNLink\" href=\"#Fn2\" id=\"#FNLinkFn2\"\u003e2\u003c/a\u003e. Additionally, the Portuguese government has seriously engaged in innovation policies, promoting initiatives like \u003cem\u003eStartup Portugal\u003c/em\u003e, \u003cem\u003e200M\u003c/em\u003e and business incubators, which has fostered the creation of several startups. Since 2016, the investment in Lisbon-based startups has grown 30% every year\u003ca class=\"FNLink\" href=\"#Fn3\" id=\"#FNLinkFn3\"\u003e3\u003c/a\u003e in a result of several successful startups and unicorns formed in Portugal. For this work, we selected eight IT startups from the \u003cem\u003eSifted\u003c/em\u003e 2020 Portugal startups list\u003ca class=\"FNLink\" href=\"#Fn4\" id=\"#FNLinkFn4\"\u003e4\u003c/a\u003e. The chosen companies are currently at different stages in their life cycle and are considered active on Twitter. The content posted by the eight startups spans five years of analysis, from 2015 to 2020, resulting in a total of 15 577 tweets.\u003c/p\u003e \u003cp\u003eThe remainder of the paper is organized as follows: after presenting the related work, the \u003cspan refid=\"Sec6\" class=\"InternalRef\"\u003emethodology\u003c/span\u003e section describes our dataset, presents the methodology and proposes a new model for the life cycle of startup business. After presenting and discussing the results of our analysis in the results section, implications are drawn. To wrap up, the conclusions section describes the main conclusions and lays the path for future work.\u003c/p\u003e\n\u003cp\u003e[1] https://www.thebalancesmb.com/top-reasons-why-your-small-business-should-use-twitter-2948523\u003c/p\u003e\n\u003cp\u003e[2] https://www.ef.com/wwen/epi/\u003c/p\u003e\n\u003cp\u003e[3] https://beportugal.com/startup-in-portugal/\u003c/p\u003e\n\u003cp\u003e[4] https://sifted.eu/portugal-startups-top-rankings/\u003c/p\u003e"},{"header":"2. Related Work","content":"\u003cp\u003eSocial media platforms play an essential role as a marketing digital tool for small businesses. Saravanakumar and Suganthalakshmi (2012) denote social media marketing (SMM) as marketing tactics that efficiently promote brands through social media platforms. However, how can we analyze social media content and extract relevant information that demonstrates this value? This section aims to answer this question by explaining the social media analysis process and its methods and results. Additionally, we describe the startups\u0026rsquo; life cycle since this constitutes the central hypothesis driving our research: the cycle of the startup\u0026rsquo;s life and evolution influences their social media activity.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. The social media analysis process and methods\u003c/h2\u003e \u003cp\u003eSocial media data has become a fundamental part of the data ecosystem and is a strategic knowledge source for decision-making (Kapoor et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). However, to infer relevant information from data, one needs to prepare and process it (Dutot and Mosconi 2016). Social media intelligence (SMI) collects and analyzes relevant data to provide data-driven support for strategic decisions. SMI works as a cycle because social media constantly changes, with new users creating new content and generating more data for analysis. The main focus of SMI applications is product/service reviews analysis (Kapoor et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). The knowledge obtained by SMI is meant to describe the present state of social media. This means that, if the objective is to predict outcomes and suggest future directions, a social media analytics (SMA) approach is deemed necessary (J. Choi et al. \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). SMA and SMI present similar phases (Zeng et al. \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2010\u003c/span\u003e), but the SMA methodology and results focus on the future while SMI concerns the present.\u003c/p\u003e \u003cp\u003eSocial media content is mainly text, and the goal for its analysis is to find relationships among data in textual documents and extract patterns to understand the themes being addressed (Jelodar et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). This can be achieved by analyzing the text\u0026rsquo;s sentiment or identifying the main topics. A topic is a list of words defined statistically to categorize the meaning of the text, and this process is termed topic modeling. In the literature, researchers address problems in the most varied fields by using topic modeling. Some of the most used methods are Latent Dirichlet Allocations (LDA) by Blei et al. (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2002\u003c/span\u003e), whose foundations are probabilistic graphical models, Latent Semantic Analysis (LSA) by Landauer et al. (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2007\u003c/span\u003e), and Non-negative Matrix Factorization (NMF) by Lee and Seung (2001), both based on linear algebra, namely, diverse forms of matrix factorization.\u003c/p\u003e \u003cp\u003eLDA or Latent Dirichelet Allocation is one of the most popular and widespread methods for identification of latent topics in a text (Blei et al. \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2002\u003c/span\u003e), and identifies the (relevant) topics by using generative probabilistic models. One of the areas where it is applied is in social media topic analysis, as observed in the works of (Saura et al. 2019; Yang and Zhang \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Yu et al. \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). While these studies focus on different problems, each one uses topic modeling as a tool for SMA. The authors of (D. Yu et al. \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) developed a novel hierarchical topic modeling technique and mined the dimension hierarchy of tweets\u0026rsquo; topics over tweeters of different countries. Saura et al. (2019) analyzed tweets with the hashtag startup (\u0026ldquo;#startup\u0026rdquo;) and its comments. The objective was to understand which topics are present in those tweets and what are the associated sentiments. Yang and Zhang (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) performed a similar analysis, where the authors combined topic modeling and sentiment analysis to mine the tweet\u0026rsquo;s text. They concluded that the LDA algorithm makes it easy to analyze an extensive set of tweets and obtain meaningful topics. Some other studies use topic modeling to explore and understand specific subjects on Twitter, like in the case of (Barry et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2018\u003c/span\u003e), which analyzes alcoholic drinks advertising, or a recent study to understand how politicians tweet about climate change by Yu et al. (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). More recent works use topic modeling methods to examine Twitter information about COVID-19. For instance, Sha et al. (\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) analyzed governmental and politicians\u0026rsquo; tweets about the pandemic situation and inferred a set of topics that describe Twitter activity over the countries under analysis. Kaila, R.P. \u0026amp; Prasad (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) and Doogan et al. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) focused on tweets bearing hashtags related to COVID-19 to understand what non-government users tweet concerning the coronavirus pandemic and its global perception. While the former studies ascertain LDA as having achieved good results in analyzing Twitter posts, they also raise limitations about the use of the LDA algorithm with Twitter data. The two most common limitations are the tweets\u0026rsquo; short text format and the need for preprocessing phase. Transforming a tweet into a document to perform a topic model might not be adequate because it has few words to extract topics. Therefore, most of the studies solve these limitations by aggregating the tweets into sets, where each collection corresponds to a document (Curiskis et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). However, some advances appear to avoid the aggregation, as the research by Xiong et al. (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2018\u003c/span\u003e), where the authors proposed a short-text topic model algorithm.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Startups and social media\u003c/h2\u003e \u003cp\u003eSocial media platforms have a global reach, are easy to access, and low cost, enabling startups to use social media as a digital marketing gateway and observe the market. A few studies investigate the potential relationships between startups and social media platforms in literature. Lugović and Ahmed (2015) found a positive correlation between startups\u0026rsquo; Twitter usage and the total investment in the source country. As previously stated, Saura et al. (2019) collected tweets presenting the hashtag startup (\u0026ldquo;#startup\u0026rdquo;). The authors aimed at relating the polarity of the tweet with the topics found within the diverse sentiments. Using the tweet\u0026rsquo;s text and its comments, classified them into positive, negative, and neutral. Then, the authors performed topic modeling for each polarity and found the correspondent topics, enabling them to understand the Twitter audience sentiment of startups related content. Ruggieri et al. (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) aimed to found patterns in successful innovative startups based on their digital platforms\u0026rsquo; activity. Their study demonstrates that startups are present on digital platforms mostly because these platforms have a cost-effective performance. The authors also conclude that a community of users/providers of services is essential for the business. Such community is fundamental for a positive impact on digital platforms, primarily in social networking web sites, since that community provides opinions, positive or negative, about products and companies. Word-of-mouth is the everyday oral communication, which creates an impression and idea about a specific subject (Keller \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2007\u003c/span\u003e) and online opinions are called electronic word-of-mouth (eWOM), as explained by Hennig-Thurau et al. (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2004\u003c/span\u003e). Social media platforms are ideal tools for eWOM. Chu and Kim (2011) describe that eWOM enables to create a large community, which allows for the increase of digital engagement with social interactions, such as comments, likes, and shares. The last two represent non-verbal activities, and when its quantities are large, might help raising a positive feeling in the social media profile in question (Wolny and Mueller 2013). Additionally, social media activities can be used to understand the online organization's reputation (Azinhaes et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. The startups\u0026rsquo; life cycle and its stages\u003c/h2\u003e \u003cp\u003eStartups are mostly defined as fast-grow innovative businesses. According to Wang et al. (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2016\u003c/span\u003e), the maturity evolution of a startup goes through two main stages: the learning stage and the growing stage. The learning stage consists of selecting a problem to be solved and the definition and evaluation of the solution. The problem represents a real issue or obstacle for a specific target, which is solved by providing a product or service: the solution. The product concept is developed in the growing stage, followed by an implementation start leading to a working prototype. In case it results, the startup obtains a functional product that later evolves into a mature product. However, Wang et al. (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2016\u003c/span\u003e) emphasizes that this is not a constant cycle, saying that a startup has to go through \u0026ldquo;multiple measure-learn loops\u0026rdquo;. The loops mean a revaluation of each step as being in the stages previously referred. Concerning startups whose main product/service is software, Nguyen-Duc et al. (2015) created a conceptual model, named the hunter-gatherer, that in fact consists of two development cycles: the \u0026ldquo;hunting\u0026rdquo; cycle consists in the idea, market, and features; the \u0026ldquo;gathering\u0026rdquo; cycle features the prototype, quality, and product. The intention is that the two cycles occur at each stage, but the dimension of the cycle differs over the startup\u0026rsquo;s cycle of life. In the learning stage, the hunting cycle is more significant, while in the growing stage, the gathering cycle becomes prominent. Nevertheless, the cycles occur at each stage side-by-side: when the company obtains a mature product, the focus changes to quality matters.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Methodology","content":"\u003cp\u003eThis research follows the SMI steps framework described by Choi et al. (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) for social media-based BI research. The process consists of four phases: \u0026ldquo;Data collection\u0026rdquo;, \u0026ldquo;Data preprocessing\u0026rdquo;, \u0026ldquo;Data analysis\u0026rdquo;, and \u0026ldquo;Validation \u0026amp; Interpretation\u0026rdquo;.\u003c/p\u003e \u003cp\u003eAccording with this framework, the initial step was the extraction of Twitter data. As previously mentioned, a particular set of startups\u0026rsquo; accounts was targeted: information technology (IT) startups founded by Portuguese or headquartered in Portugal, selling products or services based on machine learning (ML) approaches and presenting a B2B business model. Thus, our analysis centers in eight startups from the \u003cem\u003eSifted\u003c/em\u003e 2020 Portugal startups list: \u003cem\u003eAttentiveMobile\u003c/em\u003e, \u003cem\u003eCodacy\u003c/em\u003e, \u003cem\u003eDefinedCrowd\u003c/em\u003e, \u003cem\u003eFeedzai\u003c/em\u003e, \u003cem\u003eProdsmart\u003c/em\u003e, \u003cem\u003eTalkdesk\u003c/em\u003e, \u003cem\u003eUnbabel\u003c/em\u003e, and \u003cem\u003eVirtuleap\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eAfter the extraction, data was cleaned and the corpus was prepared (Data preprocessing), after which we could proceed with a topic modeling (TM) technique for the analysis (Data analysis). Finally, TM results are evaluated and interpreted (Validation \u0026amp; Interpretation). The latter step is where the topic modeling results are compared with the startups\u0026rsquo; funding rounds, creating a diachronic profile for each startup. For that, the funding rounds of each startup have also been collected from Crunchbase\u003ca class=\"FNLink\" href=\"#Fn5\" id=\"#FNLinkFn5\"\u003e5\u003c/a\u003e, and related with the startups\u0026rsquo; lifecycle phase. Our approach is illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe features that define a startup differ depending on where in the life cycle phase the company is: in the beginning, these are innovative companies with limited resources; while in the growth process, they perform an above-average rate increment in the number of customers and revenue; and finally, they have hyper-scalability and high company valuation, which characterizes a mature startup, demonstrating that startups change over their lifecycle and that the definition of a startup depends on particular phases of company\u0026rsquo;s evolution (Skala \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Thus, startups\u0026rsquo; life cycle is a complex concept and, as stated by Paschen (\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2017\u003c/span\u003e), it shows two different but connected perspectives that are fundamental for the company\u0026rsquo;s success: its maturity, regarding the stage of development of a product or service, and the funding rounds, that is, the fundamental investment attraction capability.\u003c/p\u003e \u003cp\u003eBased on the related literature, we consider that the startups\u0026rsquo; life cycle can be divided into two main perspectives. One that follows closely the concepts found in (Wang et al. \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2016\u003c/span\u003e) and regards the creation of a mature product to solve a real problem: the \u003cem\u003ematurity evolution\u003c/em\u003e. Another one concerning the startup funding rounds: the \u003cem\u003efunding rounds\u003c/em\u003e. The funding rounds are where startups open or expose their shareholder structure to third parties, usually to business angels or venture capital firms, to secure investment to allow the startup to be able to grow (Paschen \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). To illustrate a startup\u0026rsquo;s financing milestones and the startup\u0026rsquo;s evolution, we propose a life cycle model based on the previously introduced two dimensions: the \u003cem\u003efunding rounds\u003c/em\u003e and the \u003cem\u003ematurity evolution\u003c/em\u003e. We believe that the Funding and Product Evolution Model (FPEM) depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, illustrates the maturation process of a startup\u0026rsquo;s life regarding time and revenue in a typical success scenario.\u003c/p\u003e \u003cp\u003eFor the model, the names of the funding rounds dimension are based on the Crunchbase Glossary\u003ca class=\"FNLink\" href=\"#Fn6\" id=\"#FNLinkFn6\"\u003e6\u003c/a\u003e, and in the maturity evolution, the phases describe the startup\u0026rsquo;s product stages based on the work of Wang et al. (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2016\u003c/span\u003e) and Paschen (\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). The proposed model, FPEM, encompasses four key phases, named after the funding round categories: the \u003cem\u003epreseed\u003c/em\u003e phase, the \u003cem\u003eseed\u003c/em\u003e phase, the \u003cem\u003eearly\u003c/em\u003e phase, and the \u003cem\u003elate\u003c/em\u003e phase. For the creation of the model, we correlated the phases with the existing funding types since these are measurable, which is essential to be able to mark when a transition occurs. Then, we connected the product maturity evolution with each of the rounds. Therefore, a phase transition occurs with a funding round of a higher rank than the previous one, implying a scale-up for the company and a product maturity evolution. Typically, startups receive new funding when their product has evolved and created value for the company. However, every type of funding round can happen more than once throughout a company\u0026rsquo;s life. Notice that, for each phase, the association of concepts between maturity dimensions and funding rounds is relatively straightforward.\u003c/p\u003e \u003cp\u003eIn the \u003cem\u003epreseed\u003c/em\u003e phase, there is only the conceptualization of a potential and innovative solution for a concrete problem. Thus, funding is usually very limited (typically below \u003cspan\u003e$\u003c/span\u003e150K) because it finances only an idea. These funding laps are known as angel or preseed rounds and are generally used to jump-start the company, providing financial cash to build a prototype. According to Wang et al. (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2016\u003c/span\u003e), in this phase the startup is in its learning stage. Next, in the \u003cem\u003eseed\u003c/em\u003e phase, a prototype or, at least, a proof-of-concept, already exists, sustaining the seed funding which can scale up to \u003cspan\u003e$\u003c/span\u003e2M. This round is used to build a product as market ready, incorporating the novelty proposed by the startup in the previous phase. In the \u003cem\u003eearly\u003c/em\u003e phase, the company already has a functional product and is prepared for scaling in the market. In this phase, the startup evolves for the so called growing stage (Wang et al. \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). The early funding rounds, also called Series A and Series B, can have values ranging between \u003cspan\u003e$\u003c/span\u003e1M and \u003cspan\u003e$\u003c/span\u003e30M. Lastly, in the \u003cem\u003elate\u003c/em\u003e phase, a mature product is already established and the correspondent funding, also called Series C round, usually shows values that may start at \u003cspan\u003e$\u003c/span\u003e10M and with no upper limit.\u003c/p\u003e \u003cp\u003eThe above-described relations between product maturity and funding rounds that represent the proposed life cycle model are validated by the topic model approach we have obtained, whose results are discussed in Section 4. The aforementioned relations enable to relate each of the four FPEM phases with the uncovered topics extracted from the tweets posted by the startups on social media during their existence.\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e3.1. Dataset\u003c/h2\u003e \u003cp\u003eThe dataset consists of 15 577 tweets extracted from the chosen Portuguese startups\u0026rsquo; Twitter accounts. The date of extraction date January the 10th, 2021, and the data covers every tweet posted by each of the startup since its Twitter profile creation date. The Twitter API method was employed (\u0026ldquo;GET statuses/user_timeline\u0026rdquo;) to extract all the tweets posted by providing each company account\u0026rsquo;s username, through the library \u003cem\u003etweepy\u003c/em\u003e (Roesslein \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). The analysis focuses on the last five years, where the higher quantity of posts is concentrated: from January 2015 to December 2020, or equivalently, during 72 months. To accurately examine the startups\u0026rsquo; activity over time, Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the startup\u0026rsquo;s Twitter accounts\u0026rsquo; descriptions.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDataset description\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCompany Name\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFounded Date\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFirst Tweet Date\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFollowers\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNumber of Tweets\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eTweets per month\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eAttentiveMobile\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2016\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e08/02/2018\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1 115\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e695\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e20.44\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eCodacy\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2012\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e02/10/2013\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2 796\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1 640\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e22.78\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eDefinedCrowd\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2015\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e04/02/2016\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1 674\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1 258\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e21.69\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eFeedzai\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2011\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e23/10/2015\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2 630\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3 177\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e51.24\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eProdsmart\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2012\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e04/12/2012\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e897\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e211\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eTalkdesk\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2011\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e26/06/2019\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6 586\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3 211\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e178.39\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eUnbabel\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2013\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e17/11/2013\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3 510\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2 615\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e36.32\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cem\u003eVirtuleap\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2018\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e29/08/2016\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2 765\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e53.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eIt presents the company\u0026rsquo;s first tweet available date, the number of followers, the number of tweets since January 2015, and the frequency of tweets per month. The last value regards the 72 months of analysis, or the number of months since the first tweet available date if it is more recent than January 2015. Additionally, the table shows the startup founding year, collected from Crunchbase.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows each startup\u0026rsquo;s quantity of tweets distributed over our chosen time window. It is possible to see that some startups post regularly, while others present peaks with more activity. Within this context, regularly means the same temporal cadence, which is the case half of the companies in analysis, namely: \u003cem\u003eAttentiveMobile\u003c/em\u003e, \u003cem\u003eDefinedCrowd\u003c/em\u003e, \u003cem\u003eFeedzai\u003c/em\u003e, \u003cem\u003eTalkdesk\u003c/em\u003e, and \u003cem\u003eUnbabel.\u003c/em\u003e Particularly, \u003cem\u003eTalkdesk\u003c/em\u003e account presents a higher number of tweets per month.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eHowever, not every startup presents tweet\u0026rsquo;s posts since the beginning of 2015. In the cases of \u003cem\u003eAttentiveMobile\u003c/em\u003e, \u003cem\u003eDefinedCrowd\u003c/em\u003e, \u003cem\u003eFeedzai\u003c/em\u003e, and \u003cem\u003eTalkdesk\u003c/em\u003e, the date for their first tweet available is more recent (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). This may be due to the fact that the company\u0026rsquo;s foundation date is posterior or because more ancient tweets were voluntarily deleted. Namely, \u003cem\u003eFeedzai\u003c/em\u003e, and \u003cem\u003eTalkdesk\u003c/em\u003e are the \u0026lsquo;\u0026lsquo;oldest\u0026rsquo;\u0026rsquo; startups, dating from 2011, but the overall number of postings is not that high, which might suggest that they may have deleted some of their oldest tweets.\u003c/p\u003e \u003cp\u003e \u003cem\u003eCodacy\u003c/em\u003e, \u003cem\u003eProadsmart\u003c/em\u003e, and \u003cem\u003eVirtuleap\u003c/em\u003e do not post regularly, and \u003cem\u003eVirtuleap\u003c/em\u003e is the only company whose activity does not cover the 72 months of analysis time window. \u003cem\u003eCodacy\u003c/em\u003e and \u003cem\u003eVirtuleap\u003c/em\u003e present a peak in 2016 and 2017, respectively. From then on, both post with regularity but using quite fewer tweets per month. Notably, \u003cem\u003eProadsmart\u003c/em\u003e shows a considerable lesser degree of Twitter posting activity and is the only company who does not show posts in every month.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.2. Text preprocessing\u003c/h2\u003e \u003cp\u003eTo understand what the topics of the textual tweets might be, posted by the startups, we aggregated our dataset by month, resulting in a corpus (that is, the set of documents where each document has an id and the correspondent text) of 72 documents corresponding to each month in the time-scope of the analysis. Within each document, the id regards the month and year of the tweets. This corpus was then cleaned, retaining the vocabulary that accurately represents the startups\u0026rsquo; content to be transformed into a document-term matrix for model training.\u003c/p\u003e \u003cp\u003eTo assure the more adequate preprocessing of tweets, we first studied the techniques applied in literature\u0026rsquo;s similar studies, thus concluding that literature supports the need for a preprocessing phase enabling as a preparation phase for achieving coherent topics. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e presents the techniques that have been applied in the existent literature.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eLiterature preprocessing techniques usage\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePreprocessing technique\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eH. J. Choi and Park (2019)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAlash and Al-sultany (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2020\u003c/span\u003e)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDoogan et al. (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2020\u003c/span\u003e)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eHidayatullah et al. (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2018\u003c/span\u003e)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eYang and Zhang (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2018\u003c/span\u003e)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLowercase transformation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHTML tags elimination\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eURL elimination\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHashtag treatment\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRemove punctuation and digits\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRemove Stop Words\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLemmatization\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStemming\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eN-Grams\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTF-IDF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRemove extra white spaces\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRemove terms with higher frequency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRemove terms with less frequency\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eX\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe most used techniques are: URL elimination, extra white spaces elimination, and exclusion of the terms presenting higher or lesser frequency, HTML tags elimination and the usage of stop words are also commonly applied.\u003c/p\u003e \u003cp\u003eSince white spaces, URLs, and punctuation do not present information relevant towards topic\u0026rsquo;s identification, they were removed from the documents. Next, lowercase transformation and lemmatization were performed. Excluding a set of stopwords, in this case, stopwords from the \u003cem\u003eNatural Language Toolkit\u003c/em\u003e (Bird et al. \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2009\u003c/span\u003e), helps to focus the model on the relevant words that might define the text\u0026rsquo;s meaning. For this, we added the startups\u0026rsquo; names and Twitter tags, like \u0026ldquo;RT\u0026rdquo; which means that it is a retweet, to the set of stopwords. The lemmatization goal is to convert every word to a common base form, providing coherence to the set of words and, consequently, to the topics. This was achieved via \u003cem\u003eTextBlob\u003c/em\u003e library (Loria \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). \u003cem\u003eCountVectorizer\u003c/em\u003e from the Python library \u003cem\u003escikit-learn\u003c/em\u003e (Pedregosa et al. \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2011\u003c/span\u003e) enables vectorizing the text and having some preprocessing customization like the use of n-grams and exclusion of terms. The n-grams used were in a range of 1 to 2, uni to bi-grams, to gather terms that may appear together, for example the bi-gram \u0026ldquo;Machine Learning\u0026rdquo;. Then, the terms that appear less than twice were excluded to prevent possible errors and misspells. Lastly, the exclusion of terms that appear in at least 80% of the tweets. Being highly frequent terms, suggests that they are meaningless in terms of topic characterization.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.3. Topic modeling\u003c/h2\u003e \u003cp\u003eDue to its success in Twitter topic analysis related literature, the topic modeling method here employed was LDA, Latent Dirichlet Allocation (Blei et al. \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2002\u003c/span\u003e). The first step is to transform the corpus into a document-term matrix, where each term is either a word or a bigram. For that, we use the frequency of the occurrence of the term/bigram in the document\u0026rsquo;s text and apply the LDA algorithm on the resulting matrix, using the Python library \u003cem\u003egensim\u003c/em\u003e (Rehurek et al. 2011).\u003c/p\u003e \u003cp\u003eSince the number of topics must be given has input for the algorithm, we performed a coherence test for the advisable number of topics to be used in the modeling. Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e suggests that five might be the more reliable number of topics, due to its higher coherence value. Note that the coherence measure used here was \u003cem\u003ec_v\u003c/em\u003e, which is one of the options in \u003cem\u003egensim\u003c/em\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThus, the topic model created has five topics, each one characterized by the relevant terms presented in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, with all the terms showing a similar distribution within each topic.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eTopic Description\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTopic\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTerms\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFintech and ML\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003efuture, talk, fintech, banking, reality, money2020, lisbon, project, hackathon, machinelearning\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBusiness Operations\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ebusiness, cloud, opentalk2020, learn, covid19, service, solution, webinar, customer service, brand\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBank and Funding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ebank, webinar, cloud, leader, learn, read, account, report, meet, partner\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProduct/Service RD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ecloud, learn, product, read, industry, innovation, boost, service, webinar, lisbon\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ereview, codereview, analysis, learning, websummit, machinelearning, machine learning, security, staticanalysis, lisbon\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe name chosen for the first topic is \u0026ldquo;Fintech and ML\u0026rdquo; because it encapsulates \u0026ldquo;fintech\u0026rsquo;\u0026rsquo;, \u0026ldquo;machine learning\u0026rsquo;\u0026rsquo; and \u0026ldquo;banking,\u0026rdquo; as well as one event in this domain: \u0026ldquo;money2020\u0026rdquo;. The second topic is \u0026ldquo;Business Operations\u0026rdquo; since it presents terms correspondent concerns typical of the company\u0026rsquo;s operations, such as \u0026ldquo;customer service,\u0026rdquo; \u0026ldquo;brand,\u0026rdquo; \u0026ldquo;solution,\u0026rdquo; and \u0026ldquo;covid19\u0026rdquo;. Additionally, it also displays \u0026ldquo;opentalk2020\u0026rdquo;, a \u003cem\u003eTalkdesk\u003c/em\u003e\u0026rsquo;s event regarding customer service subjects. \u0026ldquo;Bank and Funding\u0026rdquo; is the third topic, supported by the terms \u0026ldquo;bank,\u0026rdquo; \u0026ldquo;leader,\u0026rdquo; \u0026ldquo;report,\u0026rdquo; and \u0026ldquo;partner,\u0026rdquo; while the fourth is \u0026ldquo;Product/Service R\u0026amp;D\u0026rdquo; sustained by terms like \u0026ldquo;innovation,\u0026rdquo; \u0026ldquo;learning,\u0026rdquo; and \u0026ldquo;boost.\u0026rdquo; Lastly, \u0026ldquo;IT\u0026rdquo; (Information Technology) is the fifth topic associated with software, like code and security, and the more significant technological event, the Websummit.\u003c/p\u003e \u003c/div\u003e\n\u003cp\u003e[5] www.crunchbase.com\u003c/p\u003e\n\u003cp\u003e[6] https://support.crunchbase.com/hc/en-us/articles/115010458467-Glossary-of-Funding-Types\u003c/p\u003e"},{"header":"4. Results And Discussion","content":"\u003cp\u003eAfter the topic model, we divided the corpus by startup and applied the model, resulting in individual analyses representing the topics\u0026rsquo; evolution over time for each one. In order to understand if there is a relation between the FPEM phases and the Twitter activity, we combined the funding rounds information. The first subsection describes the results obtained per startup.\u003c/p\u003e \u003cp\u003eAfter the individual analysis, it became clear that there are similarities between the independent analysis, so we performed another study, using all the startups\u0026rsquo; data, whose results are outlined in the second subsection.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e4.1. Topics evolution over startups life cycle\u003c/h2\u003e \u003cp\u003eThe following section regards the analysis of the Twitter activity over time for each company when combined with the startup\u0026rsquo;s funding rounds. Each figure shows the distribution of topics (in percentage), the number of tweets, and the funding rounds. To add context to the analysis, we provide, for each startup, a brief description of the company.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e represents \u003cem\u003eAttentiveMobile\u003c/em\u003e topics\u0026rsquo; evolution. \u003cem\u003eAttentiveMobile\u003c/em\u003e is a B2B company that offers a personalized mobile messaging platform. We can see that from 2015 until February 2018, there are no tweets, and social media activity starts when the startup is in its early phase, where the company already has a functional product. However, the topic \u0026ldquo;Product/Service R\u0026amp;D\u0026rdquo; is constantly present in their tweets. In 2018, \u0026ldquo;Bank and Funding\u0026rdquo; was the topic with less presence on their content, and it increased over 2019 because they may have needed new investment. Nevertheless, \u0026ldquo;Fintech and ML\u0026rdquo; and \u0026ldquo;IT\u0026rdquo; topics are always present and are half of the content posted on Twitter, which may be related to the startup product that uses machine learning techniques. Since March 2020, the topics present a stationary distribution, with a peak of tweets quantity in June 2020. This scenario with higher activity number and stable topics\u0026rsquo; distribution happens when the startup is in its late phase, where the company already has a mature product.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe evolution of \u003cem\u003eCodacy\u003c/em\u003e topics is depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. \u003cem\u003eCodacy\u003c/em\u003e is an automated code review platform. The topics distribution varies over the months, but it is clear that, in 2016, the number of tweets is significantly higher and showing two very similar peaks. During 2016, the startup is in a seed phase, which means that they should have a prototype. The startup changes to an early phase in August 2017 and, awkwardly, in September 2017, there are no tweets. The most predominant topics in its tweets are \u0026ldquo;IT\u0026rdquo;, \u0026ldquo;Fintech and ML,\u0026rdquo; and \u0026ldquo;Bank and Funding.\u0026rdquo; The first two may be related to the code review platform as it uses artificial intelligence methods, its core business, and the last may be associated with the funding needs.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eNext, \u003cem\u003eDefinedCrowd\u003c/em\u003e topics\u0026rsquo; evolution is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e. This is a company that develops artificial intelligence training data services and solutions. From 2015 until February 2016, no tweets are available, although the startup\u0026rsquo;s founding year was 2015. Until July 2018, when they received the first early round, the topics distributions fluctuate over the months, both in the number of tweets and the relative representation of topics. Once it reached its early phase, the topics started to present a more structured distribution, showing an increase of the topic \u0026ldquo;Product/Service R\u0026amp;D\u0026rdquo; in the tweets. According to the FPEM, this is a phase where, typically, companies own a fully functional product, justifying the increment of tweets related to \u0026ldquo;Product/Service R\u0026amp;D\u0026rdquo;. By the end of 2020, the graphic shows an increase in tweets per month, with two very similar peaks in July and in October.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFeedzai\u003c/em\u003e is an artificial intelligence startup, and its core business is finance risk management. \u003cem\u003eFeedzai\u003c/em\u003e tweet profile evolution can be observed in Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e. Notably, from January 2015 until November 2015, no tweets are available. From then on, Twitter activity starts with the company in an early phase with an already functional product. The topics show a stationary distribution, and the number of tweets is consistent over the months, except for peaks occurring in October 2017 and in October 2019, possibly because of an event occurring in October. Interestingly, in 2020 the topic \u0026ldquo;Bank and Funding\u0026rdquo; shows a decrease, and \u0026ldquo;Business Operations\u0026rdquo; has increased. The decrease may be since, in October 2017, the company reached the late phase, and raising more funds was no longer a priority. Alternatively, perhaps due to the COVID-19 ongoings, the company starts posting about the pandemic instead of financials-related tweets.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eProdsmart\u003c/em\u003e turns factories into digital and smart ones by employing automation mechanisms into the production and control the workflow using their software. Figure\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e represents the company\u0026rsquo;s topics evolution. The tweets\u0026rsquo; content is varied, without a visual pattern or structure, making the distribution of the topics oscillate. During 2015, April stands out with contents relating to the topic \u0026ldquo;Bank and Funding,\u0026rdquo; while in July, August, and September of the same year, the main topics in the tweets were \u0026ldquo;Product and R\u0026amp;D\u0026rdquo; and \u0026ldquo;Business Operations\u0026rdquo;. In terms of tweets quantity, there is a peak occurring in November 2016, although the number of tweets is always lower when compared with the other startups in the analysis. Since 2016, when the startup achieved the seed phase, the topics \u0026ldquo;Fintech and ML\u0026rdquo; and \u0026ldquo;IT\u0026rdquo;, representing the technology subject, start to be present in their tweets\u0026rsquo; content. Over the years, the topic \u0026ldquo;Bank and Funding\u0026rdquo; has been present, which can be explained by the company\u0026rsquo;s funding needs, since throughout the time period under analysis, \u003cem\u003eProdsmart\u003c/em\u003e did not leave the seed phase.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e represents the \u003cem\u003eTalkdesk\u003c/em\u003e topics\u0026rsquo; evolution. \u003cem\u003eTalkdesk\u003c/em\u003e is a platform to support sales teams for costumers\u0026rsquo; satisfaction and cost savings. Although it has been founded in 2011, from 2015 until July 2019, no tweets are available. However, from July 2019 forward, the number of tweets is mostly above 100/month, which suggests that they must have deleted previous posts. From this point on, \u003cem\u003eTalkdesk\u003c/em\u003e was at an early phase and reached the late phase in July 2020. Regarding Twitter\u0026rsquo;s activity, the topics are distributed very similarly over the months, with \u0026ldquo;Product R\u0026amp;D\u0026rdquo; showing the lesser number of tweets and the number of tweets showing two peaks, one in November 2019 and the other in May 2020. Since the tweets precede its entrance into a more mature phase already involving a stable product, tweeting about product development may not be between its higher priorities.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eUnbabel\u003c/em\u003e product enables companies to serve customers in their native language with a scalable translation across digital channels and Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e represents \u003cem\u003eUnbabel\u003c/em\u003e topics\u0026rsquo; evolution. Their first seed round was in March 2014. Therefore, \u003cem\u003eUnbabel\u003c/em\u003e was in a seed phase until October 2016, when it reached an early phase, followed by the late phase in September 2019. In the seed phase, the tweets\u0026rsquo; topics show an oscillatory behaviour, without a defined structure over the months. However, since the early phase, the distribution became more stationary. In September 2016, the startup shows no posts, and by September 2019, the number of tweets has decreased. Maybe because of the current late phase, they do not need to promote the product or raise more funds.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eLastly, Fig.\u0026nbsp;\u003cspan refid=\"Fig12\" class=\"InternalRef\"\u003e12\u003c/span\u003e depicts \u003cem\u003eVirtuleap\u003c/em\u003e topics\u0026rsquo; evolution. This company provides a virtual reality application that promotes brain health, with a library of games designed by neuroscientists. From 2015 to August 2016, there are no tweets available, and it is known that the company registry occurred in 2018. \u003cem\u003eVirtuleap\u003c/em\u003e achieved a seed round in February 2018. In fact, between 2018 and 2020, the company received five seed rounds. Tweets before 2018 can be found, and there is a high value peak quantity of tweets occurring in January 2017. Additionally, the topic distribution in 2017 is stationary, being the topics \u0026ldquo;Fintech and ML\u0026rdquo; and \u0026ldquo;IT\u0026rdquo; the ones with higher representation. Since 2018, the number of posts decreases throughout the year until it reaches residual values by the last quarter of 2018. Regarding the topics, the tweet content starts to be diverse and without structure, and the ones about \u0026ldquo;Product R\u0026amp;D\u0026rdquo; e \u0026ldquo;Business operations\u0026rdquo; decrease.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e4.2. Analysis of Twitter activity in life cycle phases\u003c/h2\u003e \u003cp\u003eThe previous observations suggest that the content and the number of tweets posted by the startups may differ over their FPEM life cycle phases. In fact, it is possible to see (Fig.\u0026nbsp;\u003cspan refid=\"Fig13\" class=\"InternalRef\"\u003e13\u003c/span\u003e) that the topics percentage varies in each phase.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe topic \u0026ldquo;Product R\u0026amp;D\u0026rdquo; is slightly higher in the preseed phase and the topic \u0026ldquo;Business Operations\u0026rdquo; is more eminent in the late phase. Newer companies focus more on product development, and mature ones already have a final product enabling them to post content about business concerns. The topics \u0026ldquo;Fintech and ML\u0026rdquo; and \u0026ldquo;IT\u0026rdquo; have similar distribution over the life cycle phases, being lower in preseed and seed phases and higher in early and late phases. Lastly, the topic \u0026ldquo;Bank and Funding\u0026rdquo; has a percentage higher than 20% in each phase, having its smaller value in the preseed and the higher in the seed. Concerning the phases, newer companies, in preseed and seed, post more about the topic \u0026ldquo;Bank and Funding\u0026rdquo;, demonstrating the importance of financials for them. In contrast, companies in the early and late phases have more content about the technology applied in their product, corresponding to the topics \u0026ldquo;Fintech and ML\u0026rdquo; and \u0026ldquo;IT\u0026rdquo;. Additionally, the preseed phase is the one with minor variance between the topics\u0026rsquo; percentage, showing that companies in that phase may not have a specific focus for their Twitter content.\u003c/p\u003e \u003cp\u003eFor understanding if the relative emergence of topics within tweets differs according to each of the four FPEM phases, and since we have no good reason to assume that the topics distribution follows a normal distribution, the Kruskal-Wallis test was used. This is a nonparametric method that compares the means between groups, and in this scenario, each group will be one of the four life cycle phases. For the Kruskal-Wallis test, we used the \u003cem\u003eSciPy\u003c/em\u003e (Virtanen et al. \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) library, setting the significance threshold at 0.05. The null hypothesis in this scenario is that the means on each life cycle phase are the same. If the \u003cem\u003ep-value\u003c/em\u003e is lower than the threshold, we reject the null hypothesis, meaning that the means on every life cycle are not the same. The results are presented in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, denoting the ones with a \u003cem\u003ep-value\u003c/em\u003e below the significance with (*).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eKruskal-Wallis tests results\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003ep-value\u003c/em\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTopic: Product R\u0026amp;D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(*)\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(0.00545\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTopic: IT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(*)\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(2.38\\text{E}-13\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTopic: Bank and Funding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(0.327\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTopic: Business Operations\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(*)\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(2.4E-06\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTopic: Fintech and ML\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(*)\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(8.82E-08\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of Tweets\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(*)\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(2.72E-08\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003e(*) Statistically significant (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\u0026lt; 0.05\\)\u003c/span\u003e\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe topics \u0026ldquo;Product R\u0026amp;D\u0026rdquo;, \u0026ldquo;IT\u0026rdquo;, \u0026ldquo;Business Operations\u0026rdquo; and \u0026ldquo;Fintech and ML\u0026rdquo; present a \u003cem\u003ep-value\u003c/em\u003e lower than the threshold, meaning that their means differ over the life cycle phases. \u0026ldquo;Bank and Funding\u0026rdquo; is the exception on the Kruskal-Wallis test, presenting a \u003cem\u003ep-value\u003c/em\u003e expressively higher than the significance. This might implicate that it remains more stable over the life cycle, which is consistent with the analysis of the information depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig13\" class=\"InternalRef\"\u003e13\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eThe results also prove a statistically significant relationship between the number of tweets and the startup phases. This relationship can be visualized in Fig.\u0026nbsp;\u003cspan refid=\"Fig14\" class=\"InternalRef\"\u003e14\u003c/span\u003e that shows the proportion of tweets posted per month and distributed into the life cycle phases. The graph shows that all the startups have posted more on average when traversing the preseed phase.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAdditionally, the higher variation in the preseed phase may be due to the fact that some of the startups in the analysis have been in this phase through a big part of the time window. However, posts from some other startups at a preseed phase were not available (or not included in the case where it occurred before 2015). Notoriously, once a seed phase is achieved, startups\u0026rsquo; number of posts is notably less. This may be because they have received a funding round and are now more focused on product development. Nevertheless, through the early and late phases, the number of tweets slightly increases.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig15\" class=\"InternalRef\"\u003e15\u003c/span\u003e displays the distributions of each topic to understand how, as shown by the Kruskal-Wallis test, do they differ throughout the FPEM phases. The topic \u0026ldquo;Product R\u0026amp;D\u0026rdquo; means are not equal through the life cycle phases (rejected the null hypothesis). It has higher values in the preseed phase and decreases in the subsequent ones. This change can illustrate the importance of product development in the startups\u0026rsquo; beginning and confirms the maturity stage correspondent stated in the life cycle description of FPEM. That is, startups in the preseed phase are finding a solution to a problem. The topic \u0026ldquo;Business Operations\u0026rdquo;, which means differ over the life cycle phases, has lower values in preseed and increases over the following phases. Having the opposite behaviour of \u0026ldquo;Product R\u0026amp;D\u0026rdquo; and showing that with the startup growth content about product development is exchanged by business concerns. The topics \u0026ldquo;IT\u0026rdquo; and \u0026ldquo;Fintech and ML\u0026rdquo;, related to the startups\u0026rsquo; core business in the analysis, have a similar evolution over the phases. Both topics increase until the early phase and lightly decrease in the late phase. Note that those have a statistical significance to support the means difference over the life cycle. Lastly, the topic \u0026ldquo;Bank and Funding\u0026rdquo; is the only which means do not differ over the phases, always staying around 20% value. The constant presence of this topic demonstrates the importance of fund raising and financial matters for startups and supports the fact that funding rounds are a dimension that characterizes startups.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"5. Implications","content":"\u003cp\u003eThe primary goal in this study was to understand how Twitter contents of IT startups evolve over the company\u0026rsquo;s growth. Literature shows that startups experience characteristic phases due to companies changing through their life cycle, adjusting their goals.\u003c/p\u003e \u003cp\u003eThe first contribution of this study is the conceptualization of a life cycle model. This proposal is based on two dimensions previously described in the literature: maturity evolution and funding rounds. Maturity regards the development of the product or service that the startup is selling, and the funding rounds regards capitalization through investors financing. Our proposal unites those dimensions, creating a natural flow of business evolution: the Funding and Product Evolution Model (FPEM).\u003c/p\u003e \u003cp\u003eThe second important implication of this study is the categorization of IT startups\u0026rsquo; social media activity. Understanding the Twitter content was achieved through topic modeling, leading to a well defined set of five topics describing the main subjects in the startups\u0026rsquo; tweets, which are: \u0026ldquo;Fintech and ML\u0026rdquo;; \u0026ldquo;IT\u0026rdquo;; \u0026ldquo;Business Operations\u0026rdquo;; \u0026ldquo;Product/Service R\u0026amp;D\u0026rdquo;; and \u0026ldquo;Bank and Funding\u0026rdquo;.\u003c/p\u003e \u003cp\u003eThe third implication brings light into the question of how the startup\u0026rsquo;s phases within its life cycle may affect social media usage. Our findings suggest that Twitter content produced by IT startups changes over the FPEM phases while the startups scale-up. The results outline that startups\u0026rsquo; initial posts are primarily related to product development and, in more advanced maturity phases, tweets became related to operations and business concerns. As expected, one of the topics found, \u0026ldquo;Bank and Funding\u0026rdquo;, constantly emerges in tweets over the entire life cycle, denoting financial matters are a cornerstone for startups, as it should be expected due to the particularities of these companies.\u003c/p\u003e"},{"header":"6. Conclusions And Future Work","content":"\u003cp\u003eThis study proposes a new startup\u0026rsquo;s life cycle model based on funding rounds and the companies product maturity: the Funding and Product Evolution Model - FPEM. The validity of FPEM is illustrated using an SMI cycle-based methodology to extract the main topics from eight IT startups founded by Portugueses or headquartered in Portugal. The Twitter posts were subjected to an automatic information extraction of topics to understand if the tweets\u0026rsquo; contents change while startups are scaling up. For the IT startups chosen, the tweets posted between 2015 and 2020 were subjected to a topic model analysis, adding up to a total of selected 15 577 tweets. The results were combined with the FPEM life cycle model, creating a diachronic profile for each one of the startups. It was possible to perceive that the startups\u0026rsquo; key topics are: \u0026ldquo;Fintech and ML\u0026rdquo; and \u0026ldquo;IT\u0026rdquo; which regard the startups\u0026rsquo; core business; \u0026ldquo;Business Operations\u0026rdquo; and \u0026ldquo;Product/Service R\u0026amp;D\u0026rdquo; about enterprise subjects and product development; \u0026ldquo;Bank and Funding\u0026rdquo; concerning startups\u0026rsquo; financing.\u003c/p\u003e \u003cp\u003eNevertheless, results reveal that IT startups\u0026rsquo; Twitter topics change over time according to the company\u0026rsquo;s current cycle of life. The number of tweets published also varies according to the startup phase, showing that newer and more mature IT startups post more on Twitter when compared to companies in an intermediate phase. In terms of content, the topic \u0026ldquo;Bank and Funding\u0026rdquo; is the only one of the five topics present throughout startups\u0026rsquo; life cycle, demonstrating the great importance of financial investments and capital enabling the company\u0026rsquo;s growth. On the other hand, another uncovered topic, \u0026ldquo;Product R\u0026amp;D\u0026rdquo;, is predominant during the preseed phase, showing that startups begin as product-focused companies. In contrast, the topic \u0026ldquo;Business Operations\u0026rdquo; is prevalent in the late phase, revealing that business concerns take the place of the product development content with the startup\u0026rsquo;s growth. Therefore, social media content evolves with the startups\u0026rsquo; evolution and consequently with the scaling stages.\u003c/p\u003e \u003cp\u003eLike all studies, this one is not without limitations. The first is the fact that it focused only on Portugal-based (or related) IT startups. Future research should study startups from other industries and countries to confirm if the results are similar, independently of the industry and region. This study relies only on publicly available Twitter data. Future studies should use data from other social media platforms, like LinkedIn, to understand if posted contents vary for different platforms or if complementary topics emerge. In this study, the startups were at different phases that limited the possibility of a complete startup life cycle for some. Future research could focus on studying other startups at the same phase of the FPEM for more comprehensive results. Lastly, we only validated the FPEM with the topics extracted from social media. Future work must use other data sources concerning startups to revalidate the model, like interviews with startups\u0026rsquo; founders and venture capital experts.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eWe have no known conflict of interest to disclose.\u003c/p\u003e\n\u003cp\u003eThis work was partly funded through national funds by FCT - Funda\u0026ccedil;\u0026atilde;o para a Ci\u0026ecirc;ncia e Tecnologia, I.P. under projects UIDB/04466/2020 (ISTAR), and UIDB/50021/2020 (INESC-ID).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Notes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAna Rita Peixoto\u0026nbsp; https://orcid.org/0000-0001-7618-5994\u003c/p\u003e\n\u003cp\u003eAna de Almeida \u0026nbsp;https://orcid.org/0000-0001-9519-4634\u003c/p\u003e\n\u003cp\u003eNuno Ant\u0026oacute;nio \u0026nbsp; https://orcid.org/0000-0002-4801-2487\u003c/p\u003e\n\u003cp\u003eFernando Batista \u0026nbsp;https://orcid.org/0000-0002-1075-0177\u003c/p\u003e\n\u003cp\u003eRicardo Ribeiro \u0026nbsp; https://orcid.org/0000-0002-2058-693X\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eAlash, Hayder M, and Ghaidaa A Al-sultany. 2020. \u0026ldquo;Improve Topic Modeling Algorithms Based on Twitter Hashtags Improve Topic Modeling Algorithms Based on Twitter Hashtags.\u0026rdquo;\u003c/li\u003e\n \u003cli\u003eAlotaibi, Bashayer et al. 2020. \u0026ldquo;Startup Initiative Response Analysis (SIRA) Framework for Analyzing Startup Initiatives on Twitter.\u0026rdquo; \u003cem\u003eIEEE Access\u003c/em\u003e 8: 10718\u0026ndash;30.\u003c/li\u003e\n \u003cli\u003eAzinhaes, J., F. Batista, and J. C. Ferreira. 2021. \u0026ldquo;EWOM for Public Institutions: Application to the Case of the Portuguese Army.\u0026rdquo; \u003cem\u003eSocial Network Analysis and Mining\u003c/em\u003e 11(1). https://doi.org/10.1007/s13278-021-00837-w.\u003c/li\u003e\n \u003cli\u003eBarry, Adam E., Danny Valdez, Alisa A. Padon, and Alex M. Russell. 2018. \u0026ldquo;Alcohol Advertising on Twitter\u0026mdash;A Topic Model.\u0026rdquo; \u003cem\u003eAmerican Journal of Health Education\u003c/em\u003e 49(4): 256\u0026ndash;63. https://doi.org/10.1080/19325037.2018.1473180.\u003c/li\u003e\n \u003cli\u003eBird, Steven., Ewan. Klein, and Edward. Loper. 2009. Natural language processing with Python \u003cem\u003eNatural Language Processing with Python\u003c/em\u003e. O\u0026rsquo;Reilly. https://www.oreilly.com/library/view/natural-language-processing/9780596803346/ (January 16, 2023).\u003c/li\u003e\n \u003cli\u003eBlei, David M., Andrew Y. Ng, and Michael T. Jordan. 2002. \u0026ldquo;Latent Dirichlet Allocation.\u0026rdquo; \u003cem\u003eAdvances in Neural Information Processing Systems\u003c/em\u003e 3: 993\u0026ndash;1022.\u003c/li\u003e\n \u003cli\u003eChoi, Hyeok Jun, and Cheong Hee Park. 2019. \u0026ldquo;Emerging Topic Detection in Twitter Stream Based on High Utility Pattern Mining.\u0026rdquo; \u003cem\u003eExpert Systems with Applications\u003c/em\u003e 115: 27\u0026ndash;36. https://doi.org/10.1016/j.eswa.2018.07.051.\u003c/li\u003e\n \u003cli\u003eChoi, Jaewoong et al. 2020. \u0026ldquo;Social Media Analytics and Business Intelligence Research: A Systematic Review.\u0026rdquo; \u003cem\u003eInformation Processing and Management\u003c/em\u003e 57(6): 102279. https://doi.org/10.1016/j.ipm.2020.102279.\u003c/li\u003e\n \u003cli\u003eChu, Shu-Chuan, and Yoojung Kim. 2011. \u0026ldquo;Determinants of Consumer Engagement in Electronic Word-of-Mouth (EWOM) in Social Networking Sites.\u0026rdquo; \u003cem\u003eInternational Journal of Advertising\u003c/em\u003e 30(1): 47\u0026ndash;75. https://www.tandfonline.com/doi/full/10.2501/IJA-30-1-047-075.\u003c/li\u003e\n \u003cli\u003eCuriskis, Stephan A., Barry Drake, Thomas R. Osborn, and Paul J. Kennedy. 2020. \u0026ldquo;An Evaluation of Document Clustering and Topic Modelling in Two Online Social Networks: Twitter and Reddit.\u0026rdquo; \u003cem\u003eInformation Processing and Management\u003c/em\u003e 57(2): 102034. https://doi.org/10.1016/j.ipm.2019.04.002.\u003c/li\u003e\n \u003cli\u003eCurran, Kevin, Kevin O\u0026rsquo;Hara, and Sean O\u0026rsquo;Brien. 2011. \u0026ldquo;The Role of Twitter in the World of Business.\u0026rdquo; \u003cem\u003eInternational Journal of Business Data Communications and Networking\u003c/em\u003e 7(3): 1\u0026ndash;15.\u003c/li\u003e\n \u003cli\u003eDoogan, Caitlin, Wray Buntine, Henry Linger, and Samantha Brunt. 2020. \u0026ldquo;Public Perceptions and Attitudes Toward COVID-19 Nonpharmaceutical Interventions Across Six Countries: A Topic Modeling Analysis of Twitter Data.\u0026rdquo; \u003cem\u003eJournal of medical Internet research\u003c/em\u003e 22(9): e21419.\u003c/li\u003e\n \u003cli\u003eDutot, Vincent, and Elaine Mosconi. 2016. \u0026ldquo;Social Media and Business Intelligence: Defining and Understanding Social Media Intelligence.\u0026rdquo; \u003cem\u003eJournal of Decision Systems\u003c/em\u003e 25(3): 191\u0026ndash;92.\u003c/li\u003e\n \u003cli\u003eGulati, Ranjay, and Alicia DeSantola. 2016. \u0026ldquo;Start-Ups That Last.\u0026rdquo; \u003cem\u003eHarvard Business Review\u003c/em\u003e 2016(March). https://hbr.org/2016/03/start-ups-that-last (January 16, 2023).\u003c/li\u003e\n \u003cli\u003eHennig-Thurau, Thorsten, Kevin P. Gwinner, Gianfranco Walsh, and Dwayne D. Gremler. 2004. \u0026ldquo;Electronic Word-of-Mouth via Consumer-Opinion Platforms: What Motivates Consumers to Articulate Themselves on the Internet?\u0026rdquo; \u003cem\u003eJournal of Interactive Marketing\u003c/em\u003e 18(1): 38\u0026ndash;52. https://linkinghub.elsevier.com/retrieve/pii/S1094996804700961.\u003c/li\u003e\n \u003cli\u003eHidayatullah, Ahmad Fathan et al. 2018. \u0026ldquo;Twitter Topic Modeling on Football News.\u0026rdquo; \u003cem\u003e2018 3rd International Conference on Computer and Communication Systems, ICCCS 2018\u003c/em\u003e: 94\u0026ndash;98.\u003c/li\u003e\n \u003cli\u003eJelodar, Hamed et al. 2017. \u0026ldquo;Latent Dirichlet Allocation (LDA) and Topic Modeling: Models, Applications, a Survey.\u0026rdquo; \u003cem\u003eMultimedia Tools and Applications\u003c/em\u003e 78: 183\u0026ndash;98. http://arxiv.org/abs/1711.04305.\u003c/li\u003e\n \u003cli\u003eKaila, R.P. \u0026amp; Prasad, A.V.K. 2020. \u0026ldquo;Informational Flow on Twitter - Corona Virus Outbreak \u0026ndash; Topic.\u0026rdquo; 11(3): 128\u0026ndash;34.\u003c/li\u003e\n \u003cli\u003eKapoor, Kawaljeet Kaur et al. 2018. \u0026ldquo;Advances in Social Media Research: Past, Present and Future.\u0026rdquo; \u003cem\u003eInformation Systems Frontiers\u003c/em\u003e 20(3): 531\u0026ndash;58.\u003c/li\u003e\n \u003cli\u003eKeller, Ed. 2007. \u0026ldquo;Unleashing the Power of Word of Mouth: Creating Brand Advocacy to Drive Growth.\u0026rdquo; \u003cem\u003eJournal of Advertising Research\u003c/em\u003e 47(4): 448\u0026ndash;52.\u003c/li\u003e\n \u003cli\u003eLandauer, Thomas K and McNamara, Danielle S and Dennis, Simon and Kintsch, Walter. 2007. Handbook of latent semantic analysis. \u003cem\u003eHandbook of Latent Semantic Analysis.\u003c/em\u003e Psychology Press.\u003c/li\u003e\n \u003cli\u003eLee, Daniel D, and H Sebastian Seung. 2001. \u0026ldquo;Algorithms for Non-Negative Matrix Factorization.\u0026rdquo; \u003cem\u003eAdvances in Neural Information Processing Systems\u003c/em\u003e 13. https://proceedings.neurips.cc/paper/2000/file/f9d1152547c0bde01830b7e8bd60024c-Paper.pdf (January 16, 2023).\u003c/li\u003e\n \u003cli\u003eLoria, Steven. 2020. \u0026ldquo;TextBlob: Simplified Text Processing \u0026mdash; TextBlob 0.16.0 Documentation.\u0026rdquo; https://textblob.readthedocs.io/en/dev/ (January 16, 2023).\u003c/li\u003e\n \u003cli\u003eLugović, Sergej, and Wasim Ahmed. 2015. \u0026ldquo;An Analysis of Twitter Usage Among Startups in Europe.\u0026rdquo; In , 299\u0026ndash;308. http://infoz.ffzg.hr/infuture/2015/images/papers/8-02 Lugovic, Ahmed, An Analysis of Twitter Usage Among Startups in EU.pdf.\u003c/li\u003e\n \u003cli\u003eNguyen-Duc, Anh, Pertti Sepp\u0026auml;nen, and Pekka Abrahamsson. 2015. \u0026ldquo;Hunter-Gatherer Cycle: A Conceptual Model of the Evolution of Software Startups.\u0026rdquo; \u003cem\u003eACM International Conference Proceeding Series\u003c/em\u003e 24-26-Augu(Idi): 199\u0026ndash;203.\u003c/li\u003e\n \u003cli\u003ePaschen, Jeannette. 2017. \u0026ldquo;Choose Wisely: Crowdfunding through the Stages of the Startup Life Cycle.\u0026rdquo; \u003cem\u003eBusiness Horizons\u003c/em\u003e 60(2): 179\u0026ndash;88. http://dx.doi.org/10.1016/j.bushor.2016.11.003.\u003c/li\u003e\n \u003cli\u003ePedregosa, Fabian et al. 2011. 12 Journal of Machine Learning Research \u003cem\u003eScikit-Learn: Machine Learning in Python\u003c/em\u003e. http://scikit-learn.sourceforge.net. (January 16, 2023).\u003c/li\u003e\n \u003cli\u003eRehurek, Radim and Sojka, Petr. 2011. \u0026ldquo;Gensim--Python Framework for Vector Space Modelling.\u0026rdquo; \u003cem\u003eNLP Centre, Faculty of Informatics, Masaryk University, Brno, Czech Republic\u003c/em\u003e 3.\u003c/li\u003e\n \u003cli\u003eRoesslein, Joshua. 2020. \u0026ldquo;Tweepy: Twitter for Python!\u0026rdquo; https://github.com/tweepy/tweepy (January 16, 2023).\u003c/li\u003e\n \u003cli\u003eRuggieri, Roberto et al. 2018. \u0026ldquo;The Impact of Digital Platforms on Business Models: An Empirical Investigation on Innovative Start-Ups.\u0026rdquo; \u003cem\u003eManagement and Marketing\u003c/em\u003e 13(4): 1210\u0026ndash;25.\u003c/li\u003e\n \u003cli\u003eSaravanakumar, M, and T Suganthalakshmi. 2012. \u0026ldquo;Social Media Marketing.\u0026rdquo; \u003cem\u003eLife Science Journal\u003c/em\u003e 9(4): 1097\u0026ndash;8135. http://www.lifesciencesite.comhttp//www.lifesciencesite.com.670 (January 16, 2023).\u003c/li\u003e\n \u003cli\u003eSaura, Jose Ramon, Pedro Palos-Sanchez, and Antonio Grilo.\u0026nbsp;2019. \u0026ldquo;Detecting Indicators for Startup Business Success: Sentiment Analysis Using Text Data Mining.\u0026rdquo; \u003cem\u003eSustainability (Switzerland)\u003c/em\u003e 11(3): 1\u0026ndash;14.\u003c/li\u003e\n \u003cli\u003eSha, Hao, Mohammad Al Hasan, George Mohler, and P. Jeffrey Brantingham. 2020. \u0026ldquo;Dynamic Topic Modeling of the COVID-19 Twitter Narrative among U.S. Governors and Cabinet Executives.\u0026rdquo; \u003cem\u003earXiv\u003c/em\u003e (2): 2\u0026ndash;7. http://arxiv.org/abs/2004.11692.\u003c/li\u003e\n \u003cli\u003eSkala, Agnieszka. 2019. Digital Startups in Transition Economies \u003cem\u003eDigital Startups in Transition Economies\u003c/em\u003e.\u003c/li\u003e\n \u003cli\u003eVirtanen, Pauli et al. 2020. \u0026ldquo;SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python.\u0026rdquo; \u003cem\u003eNature Methods\u003c/em\u003e 17(3): 261\u0026ndash;72. http://www.nature.com/articles/s41592-019-0686-2.\u003c/li\u003e\n \u003cli\u003eWang, Xiaofeng et al. 2016. \u0026ldquo;Key Challenges in Software Startups across Life Cycle Stages.\u0026rdquo; \u003cem\u003eLecture Notes in Business Information Processing\u003c/em\u003e 251: 169\u0026ndash;82.\u003c/li\u003e\n \u003cli\u003eWolny, Julia, and Claudia Mueller. 2013. \u0026ldquo;Analysis of Fashion Consumers\u0026rsquo; Motives to Engage in Electronic Word-of-Mouth Communication through Social Media Platforms.\u0026rdquo; \u003cem\u003eJournal of Marketing Management\u003c/em\u003e 29(5\u0026ndash;6): 562\u0026ndash;83.\u003c/li\u003e\n \u003cli\u003eXiong, Shufeng, Kuiyi Wang, Donghong Ji, and Bingkun Wang. 2018. \u0026ldquo;A Short Text Sentiment-Topic Model for Product Reviews.\u0026rdquo; \u003cem\u003eNeurocomputing\u003c/em\u003e 297: 94\u0026ndash;102. https://doi.org/10.1016/j.neucom.2018.02.034.\u003c/li\u003e\n \u003cli\u003eYang, Sidi, and Haiyi Zhang. 2018. \u0026ldquo;Text Mining of Twitter Data Using a Latent Dirichlet Allocation Topic Model and Sentiment Analysis.\u0026rdquo; \u003cem\u003eInternational Journal of Computer and Information Engineering\u003c/em\u003e 12(7): 525\u0026ndash;29.\u003c/li\u003e\n \u003cli\u003eYu, Chao et al. 2021. \u0026ldquo;Tweeting About Climate: Which Politicians Speak Up and What Do They Speak Up About?\u0026rdquo; \u003cem\u003eSocial Media + Society\u003c/em\u003e 7(3): 205630512110338. http://journals.sagepub.com/doi/10.1177/20563051211033815.\u003c/li\u003e\n \u003cli\u003eYu, Dongjin, Dengwei Xu, Dongjing Wang, and Zhiyong Ni. 2019. \u0026ldquo;Hierarchical Topic Modeling of Twitter Data for Online Analytical Processing.\u0026rdquo; \u003cem\u003eIEEE Access\u003c/em\u003e 7: 12373\u0026ndash;85.\u003c/li\u003e\n \u003cli\u003eZeng, Daniel, Hsinchun Chen, Robert Lusch, and Shu Hsing Li. 2010. \u0026ldquo;Social Media Analytics and Intelligence.\u0026rdquo; \u003cem\u003eIEEE Intelligent Systems\u003c/em\u003e 25(6): 13\u0026ndash;16.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"social-network-analysis-and-mining","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"snam","sideBox":"Learn more about [Social Network Analysis and Mining](http://link.springer.com/journal/13278)","snPcode":"13278","submissionUrl":"https://submission.nature.com/new-submission/13278/3","title":"Social Network Analysis and Mining","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Topic Modeling, Social Media, Startups, Life Cycle Model, Twitter Data","lastPublishedDoi":"10.21203/rs.3.rs-2493496/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2493496/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSocial media platforms have become powerful tools for startups, helping them find customers and raise funding. Analysing the contents posted through social media would help them make the best use of this communication and scale their business. To understand if a startup\u0026rsquo;s social media content reflects its position in its business maturation, we start by defining an adequate life cycle model for startups based on two dimensions: funding rounds and product maturity. Using Twitter as social media source of information for known Portuguese IT startups, each at their life cycle\u0026rsquo;s different phases, their tweets\u0026rsquo; data has been analyzed. Topic modeling techniques have enabled the categorization of the data according to the topics arising in the published contents, making it possible to discover that contents can be grouped into five specific topics: \u0026ldquo;Fintech and ML\u0026rdquo;, \u0026ldquo;IT\u0026rdquo;, \u0026ldquo;Business Operations\u0026rdquo;, \u0026ldquo;Product/Service R\u0026amp;D\u0026rdquo;, and \u0026ldquo;Bank and Funding\u0026rdquo;. Comparing those profiles against the startup\u0026rsquo;s life cycle to understand how contents change over time provides a diachronic profile for each company. We discovered that while some topics are prevalent in the startup\u0026rsquo;s scaling, others depend on the startup\u0026rsquo;s particular phase of the cycle, revealing that startups\u0026rsquo; Twitter social media content differs along their life cycle.\u003c/p\u003e","manuscriptTitle":"Diachronic Profile of Startup Companies through Social Media","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-01-20 16:04:41","doi":"10.21203/rs.3.rs-2493496/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-02-05T06:21:21+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-01-30T09:11:44+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"e3541d6c-82a5-4a5d-b263-a203ba3891d4","date":"2023-01-23T09:57:57+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-01-23T09:55:09+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-01-23T09:53:04+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-01-19T05:45:30+00:00","index":"","fulltext":""},{"type":"submitted","content":"Social Network Analysis and Mining","date":"2023-01-19T00:32:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"social-network-analysis-and-mining","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"snam","sideBox":"Learn more about [Social Network Analysis and Mining](http://link.springer.com/journal/13278)","snPcode":"13278","submissionUrl":"https://submission.nature.com/new-submission/13278/3","title":"Social Network Analysis and Mining","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"24895de4-e776-4203-9d1d-890304d4369a","owner":[],"postedDate":"January 20th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-10-16T20:06:13+00:00","versionOfRecord":{"articleIdentity":"rs-2493496","link":"https://doi.org/10.1007/s13278-023-01055-2","journal":{"identity":"social-network-analysis-and-mining","isVorOnly":false,"title":"Social Network Analysis and Mining"},"publishedOn":"2023-03-18 20:02:47","publishedOnDateReadable":"March 18th, 2023"},"versionCreatedAt":"2023-01-20 16:04:41","video":"","vorDoi":"10.1007/s13278-023-01055-2","vorDoiUrl":"https://doi.org/10.1007/s13278-023-01055-2","workflowStages":[]},"version":"v1","identity":"rs-2493496","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2493496","identity":"rs-2493496","version":["v1"]},"buildId":"rHA-KDH7Qsr4HCuvH75dn","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.