<?xml version="1.0" encoding="iso-8859-1" standalone="no"?>
<!DOCTYPE GmsArticle SYSTEM "http://www.egms.de/dtd/2.0.34/GmsArticle.dtd">
<GmsArticle xmlns:xlink="http://www.w3.org/1999/xlink">
  <MetaData>
    <Identifier>mibe000312</Identifier>
    <IdentifierDoi>10.3205/mibe000312</IdentifierDoi>
    <IdentifierUrn>urn:nbn:de:0183-mibe0003124</IdentifierUrn>
    <ArticleType>Research Article</ArticleType>
    <TitleGroup>
      <Title language="en">Prompting strategies for coding in AI-assisted qualitative data analysis</Title>
      <TitleTranslated language="de">Prompting-Strategien f&#252;r das Codieren in der KI-gest&#252;tzten qualitativen Datenanalyse</TitleTranslated>
    </TitleGroup>
    <CreatorList>
      <Creator>
        <PersonNames>
          <Lastname>Diaz</Lastname>
          <LastnameHeading>Diaz</LastnameHeading>
          <Firstname>Mireya</Firstname>
          <Initials>M</Initials>
          <AcademicTitleSuffix>PhD</AcademicTitleSuffix>
        </PersonNames>
        <Address>Department of Population and Quantitative Health Sciences, School of Medicine, Case Western Reserve University, Cleveland, OH 44106, United States<Affiliation>Department of Population and Quantitative Health Sciences, School of Medicine, Case Western Reserve University, Cleveland, United States</Affiliation></Address>
        <Email>mcd8&#64;case.edu</Email>
        <Creatorrole corresponding="yes" presenting="no">author</Creatorrole>
      </Creator>
    </CreatorList>
    <PublisherList>
      <Publisher>
        <Corporation>
          <Corporatename>German Medical Science GMS Publishing House</Corporatename>
        </Corporation>
        <Address>D&#252;sseldorf</Address>
      </Publisher>
    </PublisherList>
    <SubjectGroup>
      <SubjectheadingDDB>610</SubjectheadingDDB>
      <Keyword language="en">qualitative data analysis</Keyword>
      <Keyword language="en">thematic coding</Keyword>
      <Keyword language="en">artificial intelligence</Keyword>
      <Keyword language="en">ChatGPT</Keyword>
      <Keyword language="en">prompting engineering</Keyword>
      <Keyword language="en">zero-shot</Keyword>
      <Keyword language="en">few-shots</Keyword>
      <Keyword language="en">chain-of-thought</Keyword>
      <Keyword language="de">qualitative Datenanalyse</Keyword>
      <Keyword language="de">thematische Kodierung</Keyword>
      <Keyword language="de">K&#252;nstliche Intelligenz</Keyword>
      <Keyword language="de">ChatGPT</Keyword>
      <Keyword language="de">Prompt Engineering</Keyword>
      <Keyword language="de">Zero-Shot</Keyword>
      <Keyword language="de">Few-Shot</Keyword>
      <Keyword language="de">Chain-of-Thought</Keyword>
      <SectionHeading language="en">ISCB GMDS 2026</SectionHeading>
    </SubjectGroup>
    <DatePublishedList>
      <DatePublished>20260922</DatePublished>
    </DatePublishedList>
    <Language>engl</Language>
    <License license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
      <AltText language="en">This is an Open Access article distributed under the terms of the Creative Commons Attribution 4.0 License.</AltText>
      <AltText language="de">Dieser Artikel ist ein Open-Access-Artikel und steht unter den Lizenzbedingungen der Creative Commons Attribution 4.0 License (Namensnennung).</AltText>
    </License>
    <SourceGroup>
      <Journal>
        <ISSN>1860-9171</ISSN>
        <Volume>22</Volume>
        <JournalTitle>GMS Medizinische Informatik, Biometrie und Epidemiologie</JournalTitle>
        <JournalTitleAbbr>GMS Med Inform Biom Epidemiol</JournalTitleAbbr>
      </Journal>
    </SourceGroup>
    <ArticleNo>14</ArticleNo>
  </MetaData>
  <OrigData>
    <Abstract language="de" linked="yes"><Pgraph><Mark1>Einleitung:</Mark1> Die Kodierung ist ein zentraler Schritt bei vielen qualitativen Datenanalysen (QDA). K&#252;nstliche Intelligenz (KI) hat in diesen Bereich Einzug gehalten, wobei die Leistungsf&#228;higkeit stark von der verwendeten Prompting-Strategie abh&#228;ngt. Diese Arbeit untersucht die Wirksamkeit verschiedener Prompting-Strategien auf die Leistung der KI.</Pgraph><Pgraph><Mark1>Methoden:</Mark1> Die Kodierung und Themenentwicklung erfolgten anhand von Abstracts aus einer Metaanalyse von 16 Studien, die den Zusammenhang zwischen Luftverschmutzung und der Inzidenz von Pr&#228;eklampsie untersuchten. Zudem wurde ein zweiter, thematisch verwandter Datensatz analysiert, der weniger formelle Quellen enthielt. Die Kodierung wurde mit R, NVivo 14 und ChatGPT 4.1 durchgef&#252;hrt. Verschiedene Prompting-Strategien mit zunehmendem Komplexit&#228;tsgrad (Zero-Shot, Few-Shot und Chain-of-Thought &#8211; CoT) wurden hinsichtlich ihrer Genauigkeit im Vergleich zur manuellen Kodierung bewertet. Auch die direkte thematische Kodierung mittels KI wurde untersucht.</Pgraph><Pgraph><Mark1>Ergebnisse:</Mark1> Zero-Shot-Prompting (mit oder ohne CoT) lieferte bei strukturierten Texten die besten Ergebnisse. Die Anforderung einer vollst&#228;ndigen Matrix (Wort-Abstract-Zuordnung) erh&#246;hte die Sensitivit&#228;t. Die Erg&#228;nzung von CoT durch allgemeine Themeninformationen und Rollenspiele machte diese Methode zur besten Strategie f&#252;r eher umgangssprachliche Texte. Auch die direkte Abfrage thematischer Kodierungen durch die KI zeigte gute Ergebnisse. Einige Prompting-Strategien sind anf&#228;lliger f&#252;r Halluzinationen; letztere stehen im Zusammenhang mit Fehlern bei der H&#228;ufigkeitsz&#228;hlung w&#228;hrend des Kodierprozesses.</Pgraph><Pgraph><Mark1>Diskussion:</Mark1> Diese Arbeit bietet eine Leistungsbewertung eines KI-Tools f&#252;r Kodierung und Themenentwicklung unter Verwendung verschiedener Prompting-Strategien sowie einen direkten Vergleich mit g&#228;ngigen QDA-Tools auf Basis leicht verf&#252;gbarer Daten. Bei der Interaktion mit Nutzern liefern KI-Chatbots Vorschl&#228;ge f&#252;r m&#246;gliche n&#228;chste Schritte, was die Ergebnisse der Kodierung verbessern kann. KI-Tools erfordern m&#246;glicherweise Replikationen, um konsistente Ergebnisse zu gew&#228;hrleisten und ihre probabilistische Natur auszugleichen. Die Zusammenf&#252;hrung der Ergebnisse aus mindestens zwei Tools (KI und traditionelle Methoden) liefert umfassendere und besser zug&#228;ngliche Ergebnisse f&#252;r die QDA.</Pgraph></Abstract>
    <Abstract language="en" linked="yes"><Pgraph><Mark1>Introduction:</Mark1> Coding is a central step in many qualitative data analyses (QDA). Artificial intelligence (AI) has made inroads into the field with performance dependent on the prompting strategy used. This work examines the effectiveness of different prompting strategies on AI performance. </Pgraph><Pgraph><Mark1>Methods:</Mark1> Coding and theme analyses were performed on abstracts from a 16-study meta-analysis examining the association between air pollution and preeclampsia incidence. These were also performed in a second related corpus containing fewer formal sources. Coding was performed with R, NVivo 14, and ChatGPT 4.1. Several prompting strategies in increasing degree of complexity (zero-shot, few-shots, and chain-of-thought &#8211; CoT) were evaluated regarding their accuracy against manual coding. Direct thematic coding with AI was also examined. </Pgraph><Pgraph><Mark1>Results:</Mark1> Zero-shot prompting (with or without CoT) provided the best results for structured text. Requesting a full word-by-abstract matrix increases sensitivity. Supplementing CoT with general topic information and role playing made it the best strategy for more colloquial text. Directly inquiring thematic coding within AI showed good performance. Some prompting strategies are more prone to hallucinations. The latter are associated with errors in the frequency counting process while coding. </Pgraph><Pgraph><Mark1>Discussion:</Mark1> This work offers performance assessment of an AI tool for coding and theme extraction through a variety of prompting strategies combined with head-to-head comparison with popular QDA tools using readily available data. During user interaction, AI chatbots provide suggestions regarding potential next steps which may enhance results during coding. AI tools may require replication to ensure consistent findings and mitigate their probabilistic nature. Integrating results from at least two tools (AI and traditional) provides more comprehensive and accessible results for QDA.</Pgraph></Abstract>
    <TextBlock name="Introduction" linked="yes">
      <MainHeadline>Introduction</MainHeadline><Pgraph>Artificial intelligence (AI) has paved its way into almost every human activity. The text mining capabilities inherent to AI make certain processes within qualitative data analysis (QDA), such as coding, a perfect scenario where the transition in analysis assisted by AI can take place. In this setting, AI can simplify considerably the work leaving more time for the analyst to focus on interpretation and thinking.</Pgraph><Pgraph>Recent work <TextLink reference="1"></TextLink>, <TextLink reference="2"></TextLink>, <TextLink reference="3"></TextLink>, <TextLink reference="4"></TextLink>, <TextLink reference="5"></TextLink> has assessed the integration of generative AI in QDA, particularly in the role of automating the coding process. These papers show that AI performs in alignment with the human analyst; that such performance is highly dependent on the instructions or prompts given to the tool, it is susceptible to errors in quotations, code naming, and hallucination; and that we should see it as a collaborative agent rather than a substitute for the human analyst. Still, there is need for guidance on which prompts will lead to the desired results <TextLink reference="6"></TextLink> and how to best control its errors.</Pgraph><Pgraph>This paper addresses the performance of prompting in the key step of coding starting with keyword identification via word-frequency analysis, and the contribution of AI to this step compared to other traditional tools. Also, it assesses the feasibility of AI conducting thematic coding directly. These were examined in both highly structured (scientific abstracts) and less structured (interviews, opinions) text. Ultimately it seeks to fill the gap in guidance about prompting in these activities as noted by previous literature.</Pgraph></TextBlock>
    <TextBlock name="Methods" linked="yes">
      <MainHeadline>Methods</MainHeadline><Pgraph>Raw abstracts from a set of 16 studies used in a meta-analysis of air pollution and incidence of preeclampsia were used as the text for analysis <TextLink reference="7"></TextLink>. Selection of a corpus with data from public sources avoid ethical issues (privacy, informed consent) present in QDA. Abstracts were ordered by publication date to follow the availability of information as well as to determine the abstract number at which saturation was reached if present. Words identified as highly prevalent (in nine abstracts or more, equivalent to more than half of them) across the abstracts by three tools (R, NVivo and ChatGPT 4.1) were calculated. A second set of less prevalent but still prevalent words appearing in five to eight abstracts (representing the 2<Superscript>nd</Superscript> quartile of the size distribution) were ascertained as means of identifying potential themes appearing in a smaller group of abstracts. Words appearing at the top of this group (i.e., seven or eight abstracts) were selected to estimate specificity. Words related via stemming to words in the highly prevalent group were excluded from the specificity assessment, such as pollutant. Accuracy (sensitivity and specificity against manual coding assisted by R) was calculated for ChatGPT 4.1 <TextLink reference="8"></TextLink> prompting strategies.</Pgraph><SubHeadline>Coding with R</SubHeadline><Pgraph>The popular statistical language and software environment R <TextLink reference="9"></TextLink> provides a wide range of text mining functions through tidy tools using the packages tidyverse, tidytext, tibble, and dplyr. Tidy data are formatted into a column for each variable, a row for each observation, and arranged into a table. Each sentence is a unit of observation. To process the abstracts: text was tokenized into words, stop words (i.e., common words such as articles, prepositions, conjunctions, etc. that do not carry significant meaning) were removed, and frequency of occurrence of remaining words was calculated. Codes were then generated manually using the most prevalent words across abstracts (labeled as Manual Coding). Words identified by R were used as reference standard. Sensitivity was estimated as the number of true positive (i.e. words appearing in more than half of the abstracts) divided by all the highly prevalent words and transformed to a percentage. Specificity was estimated by the number of true negatives (i.e. words less prevalent) divided by the moderately prevalent words (appearing in seven or eight abstracts) and transformed to a percentage.</Pgraph><SubHeadline>Coding with NVivo</SubHeadline><Pgraph>NVivo was used to validate the word extraction and manual coding performed with the assistance of R. Two alternative word frequency options for exploration in NVivo <TextLink reference="10"></TextLink> were assessed. These are exact match with the 50 most frequent words excluding stop words and with a minimum length of two; and a second option which included the same conditions but used stemming rather than exact matching. The latter offers a comparison point for results obtained by traditional tools and AI which incorporates stemming in its processing.</Pgraph><SubHeadline>Artificial intelligence coding through ChatGPT 4.1</SubHeadline><Pgraph>ChatGPT 4.1 is the fourth generation in the large language model chatbot. Success of AI in the assigned task is highly dependent on the prompting strategy employed. This has been identified in previous work of ChatGPT within thematic analysis <TextLink reference="11"></TextLink>. Prompts are textual inputs to the generative AI tool. These allow providing the tool with questions, instructions, and&#47;or introducing the context so the tool can perform an appropriate job. The idea is to provide the tool with specificity, context, and detail to obtain better or desired outputs. A four-component framework <TextLink reference="12"></TextLink> will set the tone of the conversation between the tool and the user to make it effective. These components are: persona (define the AI&#8217;s identity, expertise); task (specify the function(s) to perform); context (provides tone and constraints); and format (defines the output&#8217;s appearance). Different prompting strategies use these components in different degrees. The main characteristics of the three types of prompting strategies that we will examine in this work are described below. For all these three strategies, the persona defined sets the depth of the response.</Pgraph><Pgraph><OrderedList><ListItem level="1" levelPosition="1" numString="1."><Mark1>Zero-shot prompting:</Mark1> it corresponds to the simplest form of prompting. This strategy lacks examples. It basically contains the task and the context.</ListItem><ListItem level="1" levelPosition="2" numString="2."><Mark1>Few-shots prompting:</Mark1> in addition to providing the task request (and context), the prompting includes examples of how the user would like the tool &#8220;to think&#8221; (including formatting cues).</ListItem><ListItem level="1" levelPosition="3" numString="3."><Mark1>Chain-of-thought (CoT) prompting:</Mark1> it was developed for complex tasks, such as numerical problems, symbolic reasoning, logic analysis <TextLink reference="13"></TextLink>. The key is to decompose the complex task in a series of steps. These are provided to the tool as a series of instructions thus the tool can think step-by-step and complete the task effectively.</ListItem></OrderedList></Pgraph><Pgraph>This paper focuses on the coding task key to many types of QDA, including thematic analysis. Coding corresponds to the second step of the six-step approach to thematic analysis <TextLink reference="14"></TextLink>. ChatGPT has been previously used with this purpose. In fact, Naeem et al. <TextLink reference="15"></TextLink> explored its efficiency compared to manual work. Two steps of this approach will be examined here: keyword selection and coding. For keyword selection AI will select the keywords from the data based on specific prompts provided to it. Within the six Rs framework to select keywords, we will emphasize repetition and richness.</Pgraph><Pgraph>In this work we evaluate the accuracy of word-frequency counting preceding coding by AI using different prompting strategies as summarized in Table 1 <ImgLink imgNo="1" imgType="table" />. This general prompting strategy was performed sequentially in three stages with increasing complexity. The most successful of the prompts in stage one (zero-shot) was selected to be the skeleton for the prompts in stage two (few-shots and role playing), and the most successful prompt from stage two as the skeleton for prompts in stage three (CoT). Also, during the interaction with the chatbot it provided information worth considering within the strategy and thus the strategy was modified adding this information within iterations (see for example Strategy 1 in Table 1 <ImgLink imgNo="1" imgType="table" /> under &#8220;added&#8221;). Each strategy is run in a new chat. For the different strategies we selected no internet access for the ChatGPT engine. This means that its web search capabilities were disabled, looking for eliciting its language synthesis functionality more than its language mimicking one. Two separate runs for each of the different prompting strategies were performed to assess consistency of findings.</Pgraph><Pgraph>Thematic coding agreement between ChatGPT 4.1 and manual coding for Strategy 8 was assessed via consensus by identifying the overlap between both sets of codes. For thematic coding agreement between replicate runs it was estimated as the percent of abstracts showing an exact match in which the codes appeared divided by all identified codes; either both replicates identified the code or both replicates did not identify the code to be counted as an agreement. </Pgraph><SubHeadline>Assessing generalizability</SubHeadline><Pgraph>A set of 12 narratives was collected from the internet covering the words air pollution and pre-eclampsia after a Google search using these terms, to assess generalizability of the findings when dealing with less formal and structured QDA data sources such as interviews, opinions, etc. Further search within the identified sites including newspapers, news channels, and YouTube was performed with the terms. The main selection criterion was that the narrative dealt with both terms. If it only dealt with one, it was discarded. The narratives identified are a mix of online newsletters commenting on journal articles <TextLink reference="16"></TextLink>, <TextLink reference="17"></TextLink>, <TextLink reference="18"></TextLink>, <TextLink reference="19"></TextLink>, <TextLink reference="20"></TextLink>, <TextLink reference="21"></TextLink>, video posts and TV news segments uploaded into YouTube <TextLink reference="22"></TextLink>, <TextLink reference="23"></TextLink>, <TextLink reference="24"></TextLink>, <TextLink reference="25"></TextLink>, <TextLink reference="26"></TextLink>. Videos were transcribed with the assistance of closed captions. Narratives were anonymized prior to analysis. After this pre-processing, the narratives were subjected to the same steps as the abstracts, that is ordering by publication or airing day, manual coding assisted with R validated with NVivo and then, ran through different prompts in ChatGPT 4.1. For the latter only the subset with best results observed for the abstracts were assessed. Accuracy (sensitivity and specificity) in selection of prevalent words (i.e., in more than half of narratives, and just close to half of the narratives respectively) were also used as the performance measure for coding prompts. Coverage of codes with respect to manual coding was used to assess direct request of thematic coding as done with abstracts. Two independent runs of the strategies were executed. </Pgraph></TextBlock>
    <TextBlock name="Results" linked="yes">
      <MainHeadline>Results</MainHeadline><SubHeadline>R text mining (word-frequency counting and manual coding) results</SubHeadline><Pgraph>R identified 27 words appearing in nine or more abstracts. Two additional terms representing numbers 95 and 10 are present in nine or more abstracts but are excluded from these 27 words given their vagueness. In our case they are part of &#8220;95&#37; confidence interval&#8221; and &#8220;PM10&#8221; (a pollutant). However, they need other terms to have meaning.</Pgraph><Pgraph>Table 2 <ImgLink imgNo="2" imgType="table" /> illustrates that information in these abstracts can be categorized into nine codes. All these codes can be identified (using words appearing in more than half of the abstracts, and seven of them are well represented (Methods and Features are poorly represented). All the codes can be well identified using words appearing in more than a quarter of the abstracts (i.e., five or more abstracts).  </Pgraph><SubHeadline>NVivo 14 results</SubHeadline><Pgraph>Only two words were different in frequency between R text mining code and NVivo 14 using exact matching of words. Five words detected as prevalent in nine or more abstracts by R did not make the cut among the 50 more prevalent words within NVivo. Using the stemming feature in NVivo 14, identifies 12 words in which more instances of the word are encountered by NVivo 14 than by R, which continues doing exact matching. In these 12 instances an average of 8.8 more words in overall is identified by NVivo 14. In general, NVivo 14 registers where in the text the words appear. However, these should be searched one by one rather than appearing in a summary table as it is possible with both R and ChatGPT. Another matching strategy available in NVivo 14 is matching with generalizations. Although not used in this work, this strategy applies relationships among words that can help identifying potential codes.  </Pgraph><SubHeadline>ChatGPT 4.1 results</SubHeadline><Pgraph>Figure 1 <ImgLink imgNo="1" imgType="figure" /> shows a receiving operating characteristic (ROC) curve summarizing the accuracy results obtained with the different prompting strategies to ChatGPT 4.1 for the first run (gray circles), the second run (white squares), and their average (dark gray diamonds). Depending on the strategy, ChatGPT 4.1 identifies between one quarter to three quarters of the words in the set of 27 highly prevalent words for the first run, but only 44&#37; in the second run. Strategies S1e (zero shot with full word by abstract matrix) and S4b (zero shot chain-of-thought with full word by abstract matrix) are the ones with the greatest sensitivity for the first run. When both runs were averaged, sensitivity was below 70&#37; for all strategies. </Pgraph><Pgraph>There are words counted in more abstracts than in which they effectively appear, an average of two abstracts more. This corresponds to the phenomenon known as hallucination, and the different prompting strategies aimed at improving the recognition of the words by the engine trying to avoid these hallucinations, not necessarily successfully in all the strategies though. Hallucinations also manifest in identifying additional words that are not as prevalent &#8211; not at least at the level of nine or more abstracts. These correspond to false positives under the condition of highly prevalent words. Twelve words were under this condition and were selected to assess the specificity of the tool with CoT (S5) showing better performance in this measure.</Pgraph><Pgraph>Table 3 <ImgLink imgNo="3" imgType="table" /> illustrates the coding generated via manual coding with information from the keyword extracted in R and in NVivo. It also shows the coding generated by ChatGPT 4.1 when asked directly to perform thematic coding without the process of keyword extraction (Strategy 8). Both processes generate nine codes which explain the information contained in the 16 abstracts. Although they contain the same number of codes, they do not necessarily match on a one-to-one fashion. For example, &#8220;Data source&#8221; and &#8220;Measures,&#8221; and to a lesser extent &#8220;Host,&#8221; are codes identified from manual coding while these are not selected by ChatGPT 4.1. What manual coding identified as &#8220;Methods,&#8221; ChatGPT was more specific as to whether these methods corresponded to measure the exposure or to perform the analysis, a valid distinction. Likewise, the two manual codes &#8220;Features&#8221; and &#8220;Findings,&#8221; ChatGPT 4.1 differentiates them into &#8220;Effect modification&#8221; and &#8220;Exposure timing,&#8221; and into &#8220;Association strength&#8221; and &#8220;Confounding adjustment&#8221; respectively.</Pgraph><Pgraph>Table 4 <ImgLink imgNo="4" imgType="table" /> summarizes the themes identified by ChatGPT 4.1 using strategies 8a and 8b for the first run (Total and S8a), for the second run (Total-2 and S8a-2), and the percent agreement in terms of abstracts in which themes are identified or not by both runs (&#37;Agree). When we alter Strategy 8a by adding one abstract at a time (Strategy 8b) rather than providing them all at once, we can identify how the different codes emerge sequentially. We can observe in Table 4 <ImgLink imgNo="4" imgType="table" /> that codes identified in Strategy 8a correspond to those which appear in seven or more abstracts except for the notion of &#8220;air pollution&#8221; which although it appears in the sixteen abstracts it is not indicated in a general context, but it is represented by the specific pollutants studied. Table 4 <ImgLink imgNo="4" imgType="table" /> also indicates that saturation has not been reached yet because two codes emerge from the four last abstracts. These are &#8220;Joint effects and multiple exposures,&#8221; and &#8220;Built and natural environment.&#8221; We could embed them within a more general code &#8220;Environmental exposure&#8221; or simply &#8220;Exposure&#8221; as suggested by manual coding. Another issue that Strategy 8b reveals is a certain degree of variability in the results obtained via ChatGPT 4.1. For example, &#8220;Confounding adjustment&#8221; is identified as a code in 12 abstracts by Strategy 8a but only in seven by Strategy 8b. Likewise &#8220;Association strength&#8221; is identified in 16 abstracts by Strategy 8a but only in 12 in Strategy 8b under &#8220;Magnitude &#38; relevance&#47;consistency of effects&#8221;. </Pgraph><Pgraph>Agreement between the two runs for strategy S8b varied between 44&#37; and 100&#37;. &#8220;Pre-eclampsia as main outcomes&#8221; appears as a main theme separately from other adverse outcomes in the second run. This was much less marked in the first run, corresponding to the lowest level of agreement at 44&#37;. Another theme that differed somewhat between the two sets of runs is that of &#8220;Quantification of risk&#8221;. In the second run, it gave priority to a &#8220;Dose-response&#47;exposure gradient&#8221; while it was &#8220;Quantification of risk&#8221; per se in the first run.</Pgraph><Pgraph>For the narratives, the subset with best results observed for the abstracts (S1e, S4a, S4b, S5, S7b) were assessed. Thirteen words appeared in seven or more abstracts and formed the set to estimate sensitivity while 10 words appeared in five or six abstracts and formed the set to estimate specificity. Sensitivity ranged between 46&#37; to 77&#37; with S7b showing the greatest sensitivity, and S5 the second best for the first run (Figure 2 <ImgLink imgNo="2" imgType="figure" />). It ranged between 77&#37; and 85&#37; for the second run. Specificity ranged between 80&#37; and 90&#37; with S7b exhibiting the smallest specificity and S5 and S4b the highest one for the first run. It ranged between 50&#37; and 100&#37; for the second run. These results contrast with those from the abstracts. For the abstracts S4b offered the best compromise between sensitivity (though poor) and specificity between the two runs, while for the narratives S7b was consistently the best performing prompting strategy followed by S5, and then S4b.</Pgraph><Pgraph>Five themes (Health condition, Exposure, Findings, Host, and Pollutant separate from Exposure) appeared using the most prevalent words approach in manual coding. A sixth theme (Methods) appeared using less prevalent words. All these themes are shared with the abstracts. When prompted directly for thematic coding, ChatGPT 4.1 also identified these six themes appearing in a more elaborate fashion and with sub-themes. In terms of agreement between the two runs, 15 codes emerged, with 8 codes overlapping (53&#37;), one specific to the first run (Research Gaps and Limitations), and six codes specific to the second run, including &#8220;Noise Pollution&#8221;, &#8220;Epidemiologic Evidence&#8221;, and &#8220;Causality Caution&#8221;.</Pgraph></TextBlock>
    <TextBlock name="Discussion" linked="yes">
      <MainHeadline>Discussion</MainHeadline><Pgraph>Data coding, a key step in many QDA, is a perfect activity to experiment and become familiar with AI tools within the analytic classroom. We examined its performance and compared it with more traditional tools such as text analysis through the popular R and NVivo. Selecting abstracts from meta-analyses with a moderate number of studies allows examining these tools and learning the process of data coding without confronting issues of data privacy and informed consent.</Pgraph><Pgraph>Data coding with R and NVivo 14 was done through a process of identifying prevalent words via frequency-counting across the abstracts followed by building codes based on the meaning and relationship of these prevalent words. The process of frequency counting is very accurate in both tools. The same process was followed using ChatGPT 4.1. In this case the accuracy and thus identification of prevalent words depends on the prompting strategy used. Three different sets of strategies were assessed with increasing level of complexity from zero-shot prompting, to few-shots prompting, to the more elaborate CoT prompting. Zero shot (conventional or through a CoT) accompanied by a display of a full word by abstract matrix provided the best results for structured texts. Although these strategies presented a better performance, still frequency-counting and identification of words complying with a specific prevalence threshold by ChatGPT 4.1 are deficient. This process of finding highly prevalent words for coding and thematic analysis provides a practical way of finding shared themes contained in a corpus. However, if we want to find all themes, we need to be more flexible and consider less prevalent words too. Highly prevalent words are consonant with the concept of repetition for identifying keywords. However, it does not equal thematic relevance. Important themes may be discussed infrequently, as indeed it was highlighted by less prevalent words and how they defined codes not detected by the highly prevalent words. Another study <TextLink reference="27"></TextLink> noticed that ChatGPT (version 4.0) failed to detect low-frequency codes, or that it suggests themes without subthemes. In our case, neither situation arose; rare codes were identified, as well as themes and subthemes. The limitation was in fact introduced by design, that is by specifying a minimum number of abstracts in which these words should appear.</Pgraph><Pgraph>To assess the generalizability of results in terms of prompting strategies performance, less structured text was sought whose abstracts addressed the same overarching topic. Indeed, narratives included a large subset of themes contained in the abstracts in an informal tone. In the more informal sources, we ascertain that some of the previous findings hold, i.e., zero shot still providing good performance with or without CoT. However, adding general topic information and role playing to CoT provided the best result with both good sensitivity and specificity. The latter allows us to obtain most of the words that will assist in identifying codes. </Pgraph><Pgraph>A somewhat surprising finding is that CoT does not necessarily work as the best strategy in all settings. It did not for structured text, however it resulted in the best strategy when assisted with overall topic information and role playing for less structured text. Likely structured text requires less sophisticated prompting to uncover its content, in contrast to text including a more colloquial tone. However, the overall goal of identifying highly prevalent words includes multi-step reasoning, a scenario in which CoT excels and which is the reason why CoT was chosen. Others have identified situations in which CoT does not perform as well as expected. It can happen when </Pgraph><Pgraph><OrderedList><ListItem level="1" levelPosition="1" numString="1.">too many steps are involved causing problems in one or more intermediate steps (known as reasoning inflation which includes self-doubt and verification, back-tracking and retries, or repetition of previous steps <TextLink reference="28"></TextLink>); </ListItem><ListItem level="1" levelPosition="2" numString="2.">certain prompts are ignored in favor of memorized facts (known as prior knowledge bias produced by prior information that prevails over new information even when introduced in the prompt <TextLink reference="29"></TextLink>); or </ListItem><ListItem level="1" levelPosition="3" numString="3.">novel reasoning models already incorporate this stepping process internally and thus resist receiving it through the prompts.</ListItem></OrderedList></Pgraph><Pgraph>Hallucination, a concept related to validity and trustworthiness, in our case and likely in the task of coding from keyword extraction arises from errors in the frequency counting process, as evidenced by the full word by abstract matrix. It is not clear if this arithmetic failure is in the counting per se, or in keeping track which words are observed in which abstract since both issues can be appreciated in those matrices. Using ChatGPT 4.1 to generate the codes via the frequency count process would have missed a third of the codes while another third would have been more detailed. The zero-shot prompting strategies with full word matrix serve as guidance for practitioners as what can be used for coding, while for learners exploring all strategies should be part of the learning process of both coding in QDA and use of AI.</Pgraph><Pgraph>We noted that thematic saturation was not reached after 16 abstracts, a number close to empirical saturation thresholds in many QDA. However, in our example the new codes appearing from the later abstracts (more recent in time) reflect the lack of saturation in the field itself. Research related to joint effects of pollutants and multiple exposures, as well as in depth assessment of the built environment is more recent in time in comparison to research dealing with the other codes identified.</Pgraph><Pgraph>Asking ChatGPT 4.1 to generate the data codes directly without using a word frequency counting strategy led to the best results for coding itself. It identified most codes we extracted using a manual approach. In our specific example nine codes summarize well the content of these abstracts and the underlying themes developed. The ability of ChatGPT to perform thematic analysis has been evaluated by several authors <TextLink reference="4"></TextLink>, <TextLink reference="5"></TextLink>, <TextLink reference="11"></TextLink>, <TextLink reference="30"></TextLink>, <TextLink reference="31"></TextLink> who also assessed single prompting strategies and how these could augment or automate coding and thematic analysis. Most of these studies evaluated either zero-shot or few-shot prompting in earlier versions of ChatGPT (3.5 and 4.0). These authors identified similar advantages (speed, reasonable coding, and theme identification) and limitations (hallucinations) as the present work did. Only one study <TextLink reference="32"></TextLink> has assessed the same three prompting strategies evaluated here (still in a previous ChatGPT version) in a corpus of 14 texts of various sources (magazines, blog posts, news articles, transcript) and sizes (252 words to 21.6 thousand words). These authors are more critical of their findings in comparison to the current and previous works. They also had greater expectations of the role of AI, which is fully automation of the coding and thematic analysis processes. They needed eight cycles of conversations to start observing results that conformed to their expectations. They indicated that demand for a manual oversight would negate any ascribed efficiency gains. They also indicated that these previous studies which obtained comparable results drew opposite conclusions. So, it seems that the level of satisfaction with AI&#8217;s performance for QDA will depend on the users&#8217; expectations of AI role, i.e., AI as an assistant vs. an independent analyst.</Pgraph><Pgraph>Combining the codes derived manually and those generated by ChatGPT 4.1 provides a more comprehensive and detailed structuring of the information contained in the corpus analyzed. This indicates that likely using a couple of these tools rather than just one tool can assist analysts to a more successful QDA. Likewise, when using AI tools, it will be advisable to run at least a couple of replicates of the same process to determine consistency in the findings given the random nature of its underlying learning process. We observed in both structured and less structured corpora, that although there is at least a 50&#37; consensus between two runs, independent codes emerge in individual runs and thus running at least a couple of them increases the chance to tap into those non-overlapping codes and themes. This issue relates to reliability. This is necessary in case that the parameter temperature that controls the level of randomness in the chatbot is not accessible to the user as frequently occurs. Additionally, performing independent runs would mimic the process of different coders analyzing the data.</Pgraph><Pgraph>It is important to warn readers of two issues: one relates to the scope of the findings, the other to comparisons of results across studies involving AI. The findings should be considered in the context of the two related corpora examined and be cautious when extrapolating to other corpora. Also, it is very important to consider the version of the AI tool &#8211; given that from one generation to the next, and even within the same generation with few modifications (e.g. from ChatGPT 4 to ChatGPT 4.1) &#8211; may lead to substantial enhancements to the tools that may affect positively QDA processes and thus issues raised with a previous version may be solved by those modifications of a newer version.</Pgraph></TextBlock>
    <TextBlock name="Conclusion" linked="yes">
      <MainHeadline>Conclusion</MainHeadline><Pgraph>AI chatbots simplify the process of coding within QDA. Either, simple zero shot strategies for structured text, chain-of-thought with supplemental general topic information and role playing for less structured text, or a direct request of thematic coding combined with a classical QDA tool allow the analyst to derive a comprehensive set of codes. Users of both QDA and AI should be exposed to the whole process of prompting to gain a more comprehensive and rewarding learning experience. </Pgraph></TextBlock>
    <TextBlock name="Note" linked="yes">
      <MainHeadline>Note</MainHeadline><SubHeadline>Competing interests</SubHeadline><Pgraph>The author declares that she has no competing interests.</Pgraph></TextBlock>
    <References linked="yes">
      <Reference refNo="1">
        <RefAuthor>Gao J</RefAuthor>
        <RefAuthor>Guo Y</RefAuthor>
        <RefAuthor>Lim G</RefAuthor>
        <RefAuthor></RefAuthor>
        <RefTitle>CollabCoder: A lower-barrier, rigorous workflow for inductive collaborative qualitative analysis with large language models</RefTitle>
        <RefYear>2024</RefYear>
        <RefBookTitle>CHI &#8217;24: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems; 2024 May 11-16; Honolulu, HI, USA</RefBookTitle>
        <RefPage></RefPage>
        <RefTotal>Gao J, Guo Y, Lim G, et al. CollabCoder: A lower-barrier, rigorous workflow for inductive collaborative qualitative analysis with large language models. In: Floyd Mueller F, Kyburz P, editors. CHI &#8217;24: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems; 2024 May 11-16; Honolulu, HI, USA. New York, NY: Association for Computing Machinery; 2024. Article No. 11. DOI: 10.1145&#47;3613904.3642002</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1145&#47;3613904.3642002</RefLink>
      </Reference>
      <Reference refNo="2">
        <RefAuthor>Morgan DL</RefAuthor>
        <RefTitle>Exploring the use of artificial intelligence for qualitative data analysis: The case of ChatGPT</RefTitle>
        <RefYear>2023</RefYear>
        <RefJournal>International Journal of Qualitative Methods</RefJournal>
        <RefPage>1-10</RefPage>
        <RefTotal>Morgan DL. Exploring the use of artificial intelligence for qualitative data analysis: The case of ChatGPT. International Journal of Qualitative Methods. 2023 Oct 30;22:1-10. 
DOI: 10.1177&#47;16094069231211248</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1177&#47;16094069231211248</RefLink>
      </Reference>
      <Reference refNo="3">
        <RefAuthor>Xiao Z</RefAuthor>
        <RefAuthor>Yuan X</RefAuthor>
        <RefAuthor>Liao Q</RefAuthor>
        <RefAuthor>Abdelghani R</RefAuthor>
        <RefAuthor>Ouedeyer PY</RefAuthor>
        <RefTitle>Supporting qualitative analysis with large language models: Combining codebook with GPT-3 with deductive coding</RefTitle>
        <RefYear></RefYear>
        <RefBookTitle>Companion Proceedings of the 28th International Conference on Intelligent User Interfaces; 2023 Mar 27</RefBookTitle>
        <RefPage>75-78</RefPage>
        <RefTotal>Xiao Z, Yuan X, Liao Q, Abdelghani R, Ouedeyer PY. Supporting qualitative analysis with large language models: Combining codebook with GPT-3 with deductive coding. In: Companion Proceedings of the 28th International Conference on Intelligent User Interfaces; 2023 Mar 27. p. 75-78. 
DOI: 10.1145&#47;3581754.3584136</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1145&#47;3581754.3584136</RefLink>
      </Reference>
      <Reference refNo="4">
        <RefAuthor>De Paoli S</RefAuthor>
        <RefTitle>Performing an inductive thematic analysis of semi-structured interviews with a large language model: An exploration and provocation on the limits of the approach</RefTitle>
        <RefYear>2024</RefYear>
        <RefJournal>Social Science Computer Review</RefJournal>
        <RefPage>997-1019</RefPage>
        <RefTotal>De Paoli S. Performing an inductive thematic analysis of semi-structured interviews with a large language model: An exploration and provocation on the limits of the approach. Social Science Computer Review. 2024 Aug;42(4):997-1019. 
DOI: 10.1177&#47;08944393231220483</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1177&#47;08944393231220483</RefLink>
      </Reference>
      <Reference refNo="5">
        <RefAuthor>Turobov A</RefAuthor>
        <RefAuthor>Coyle D</RefAuthor>
        <RefAuthor>Harding V</RefAuthor>
        <RefTitle></RefTitle>
        <RefYear>2024</RefYear>
        <RefBookTitle>Using ChatGPT for thematic analysis. Working Paper</RefBookTitle>
        <RefPage></RefPage>
        <RefTotal>Turobov A, Coyle D, Harding V. Using ChatGPT for thematic analysis. Working Paper. Bennet Institute for Public Policy, University of Cambridge; 2024.</RefTotal>
      </Reference>
      <Reference refNo="6">
        <RefAuthor>Nguyen-Trung K</RefAuthor>
        <RefTitle>ChatGPT in thematic analysis: Can AI become a research assistant in qualitative research&#63;</RefTitle>
        <RefYear>2025</RefYear>
        <RefJournal>Quality and Quantity</RefJournal>
        <RefPage>4945-4978</RefPage>
        <RefTotal>Nguyen-Trung K. ChatGPT in thematic analysis: Can AI become a research assistant in qualitative research&#63; Quality and Quantity. 2025 Jun 2;59:4945-4978. DOI: 10.1007&#47;s11135-025-02165-z</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1007&#47;s11135-025-02165-z</RefLink>
      </Reference>
      <Reference refNo="7">
        <RefAuthor>Bai W</RefAuthor>
        <RefAuthor>Li Y</RefAuthor>
        <RefAuthor>Niu Y</RefAuthor>
        <RefAuthor>Ding Y</RefAuthor>
        <RefAuthor>Yu X</RefAuthor>
        <RefAuthor>Zhu B</RefAuthor>
        <RefAuthor>Duan R</RefAuthor>
        <RefAuthor>Duan H</RefAuthor>
        <RefAuthor>Kou C</RefAuthor>
        <RefAuthor>Li Y</RefAuthor>
        <RefAuthor>Sun Z</RefAuthor>
        <RefTitle>Association between ambient air pollution and pregnancy complications: A systematic review and meta-analysis of cohort studies</RefTitle>
        <RefYear>2020</RefYear>
        <RefJournal>Environ Res</RefJournal>
        <RefPage>109471</RefPage>
        <RefTotal>Bai W, Li Y, Niu Y, Ding Y, Yu X, Zhu B, Duan R, Duan H, Kou C, Li Y, Sun Z. Association between ambient air pollution and pregnancy complications: A systematic review and meta-analysis of cohort studies. Environ Res. 2020 Jun;185:109471. 
DOI: 10.1016&#47;j.envres.2020.109471</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1016&#47;j.envres.2020.109471</RefLink>
      </Reference>
      <Reference refNo="8">
        <RefAuthor>OpenAI</RefAuthor>
        <RefTitle></RefTitle>
        <RefYear>2023</RefYear>
        <RefBookTitle>ChatGPT (version 4.1) &#91;Large language model</RefBookTitle>
        <RefPage></RefPage>
        <RefTotal>OpenAI. ChatGPT (version 4.1) &#91;Large language model&#93;. 2023. Available from: https:&#47;&#47;chat.openai.com&#47;chat</RefTotal>
        <RefLink>https:&#47;&#47;chat.openai.com&#47;chat</RefLink>
      </Reference>
      <Reference refNo="9">
        <RefAuthor>R Core Team</RefAuthor>
        <RefTitle></RefTitle>
        <RefYear>2021</RefYear>
        <RefBookTitle>R: A language and environment for statistical computing</RefBookTitle>
        <RefPage></RefPage>
        <RefTotal>R Core Team. R: A language and environment for statistical computing. Vienna, Austria: R Foundation for Statistical Computing; 2021. Available from: https:&#47;&#47;www.R-project.org</RefTotal>
        <RefLink>https:&#47;&#47;www.R-project.org</RefLink>
      </Reference>
      <Reference refNo="10">
        <RefAuthor>Lumivero</RefAuthor>
        <RefTitle></RefTitle>
        <RefYear>2023</RefYear>
        <RefBookTitle>NVivo (Version 14) &#91;Computer software&#93;</RefBookTitle>
        <RefPage></RefPage>
        <RefTotal>Lumivero. NVivo (Version 14) &#91;Computer software&#93;. 2023. Available from: https:&#47;&#47;lumivero.com&#47;products&#47;vivo</RefTotal>
        <RefLink>https:&#47;&#47;lumivero.com&#47;products&#47;vivo</RefLink>
      </Reference>
      <Reference refNo="11">
        <RefAuthor>Lee VV</RefAuthor>
        <RefAuthor>van der Lubbe SCC</RefAuthor>
        <RefAuthor>Goh LH</RefAuthor>
        <RefAuthor>Valderas JM</RefAuthor>
        <RefTitle>Harnessing ChatGPT for Thematic Analysis: Are We Ready&#63;</RefTitle>
        <RefYear>2024</RefYear>
        <RefJournal>J Med Internet Res</RefJournal>
        <RefPage>e54974</RefPage>
        <RefTotal>Lee VV, van der Lubbe SCC, Goh LH, Valderas JM. Harnessing ChatGPT for Thematic Analysis: Are We Ready&#63; J Med Internet Res. 2024 May 31;26:e54974. DOI: 10.2196&#47;54974</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.2196&#47;54974</RefLink>
      </Reference>
      <Reference refNo="12">
        <RefAuthor>Case Western Reserve University</RefAuthor>
        <RefTitle></RefTitle>
        <RefYear>2026</RefYear>
        <RefBookTitle>AI Prompt Engineering 101: Foundational Skills. Training Workshop</RefBookTitle>
        <RefPage></RefPage>
        <RefTotal>Case Western Reserve University. AI Prompt Engineering 101: Foundational Skills. Training Workshop. 2026 Mar 3. Available from: https:&#47;&#47;case.edu&#47;hr&#47;professional-development&#47;live-training-demand-courses&#47;ai-prompt-engineering</RefTotal>
        <RefLink>https:&#47;&#47;case.edu&#47;hr&#47;professional-development&#47;live-training-demand-courses&#47;ai-prompt-engineering</RefLink>
      </Reference>
      <Reference refNo="13">
        <RefAuthor>Wei J</RefAuthor>
        <RefAuthor>Wang X</RefAuthor>
        <RefAuthor>Schuurmans D</RefAuthor>
        <RefAuthor>Bosma M</RefAuthor>
        <RefAuthor>Ichter B</RefAuthor>
        <RefAuthor>Xia F</RefAuthor>
        <RefAuthor>Ed Chi E</RefAuthor>
        <RefAuthor>Le Q</RefAuthor>
        <RefAuthor>Zhou D</RefAuthor>
        <RefTitle>Chain-of-thought prompting elicits reasoning in large language models</RefTitle>
        <RefYear></RefYear>
        <RefBookTitle>NIPS&#8217;22: Proceedings of the 36th Conference on Neural Information Processing Systems; 2022 Nov 28 - Dec 9; New Orleans (LA), USA</RefBookTitle>
        <RefPage>24824-24837</RefPage>
        <RefTotal>Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Ed Chi E, Le Q, Zhou D. Chain-of-thought prompting elicits reasoning in large language models. In: NIPS&#8217;22: Proceedings of the 36th Conference on Neural Information Processing Systems; 2022 Nov 28 - Dec 9; New Orleans (LA), USA. (Advances in Neural Information Processing Systems; 35). p. 24824-24837. 
DOI: 10.52202&#47;068431-1800</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.52202&#47;068431-1800</RefLink>
      </Reference>
      <Reference refNo="14">
        <RefAuthor>Braun V</RefAuthor>
        <RefAuthor>Clarke V</RefAuthor>
        <RefTitle>Using thematic analysis in psychology</RefTitle>
        <RefYear>2006</RefYear>
        <RefJournal>Qualitative Research in Psychology</RefJournal>
        <RefPage>77-101</RefPage>
        <RefTotal>Braun V, Clarke V. Using thematic analysis in psychology. Qualitative Research in Psychology. 2006;3(2):77-101. 
DOI: 10.1191&#47;1478088706qp0630a</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1191&#47;1478088706qp0630a</RefLink>
      </Reference>
      <Reference refNo="15">
        <RefAuthor>Naeem M</RefAuthor>
        <RefAuthor>Smith T</RefAuthor>
        <RefAuthor>Thomas L</RefAuthor>
        <RefTitle>Thematic analysis and artificial intelligence: A step-by-step process for using ChatGPT in thematic analysis</RefTitle>
        <RefYear>2025</RefYear>
        <RefJournal>International Journal of Qualitative Methods</RefJournal>
        <RefPage>1-18</RefPage>
        <RefTotal>Naeem M, Smith T, Thomas L. Thematic analysis and artificial intelligence: A step-by-step process for using ChatGPT in thematic analysis. International Journal of Qualitative Methods. 2025;24:1-18. DOI: 10.1177&#47;16094069251333886</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1177&#47;16094069251333886</RefLink>
      </Reference>
      <Reference refNo="16">
        <RefAuthor>Anonym</RefAuthor>
        <RefTitle>One in 20 cases of pre-eclampsia may be linked to air pollutant</RefTitle>
        <RefYear>2013</RefYear>
        <RefJournal>Science Daily</RefJournal>
        <RefPage></RefPage>
        <RefTotal>One in 20 cases of pre-eclampsia may be linked to air pollutant. Science Daily; 2013 Feb 7. Available from: https:&#47;&#47;www.sciencedaily.com&#47;releases&#47;2013&#47;02&#47;130206185852.htm</RefTotal>
        <RefLink>https:&#47;&#47;www.sciencedaily.com&#47;releases&#47;2013&#47;02&#47;130206185852.htm</RefLink>
      </Reference>
      <Reference refNo="17">
        <RefAuthor>Sinpetry L</RefAuthor>
        <RefTitle>Air Pollution Linked to 1 in 20 Cases of Pre-Eclampsia</RefTitle>
        <RefYear>2013</RefYear>
        <RefJournal>Softpedia</RefJournal>
        <RefPage></RefPage>
        <RefTotal>Sinpetry L. Air Pollution Linked to 1 in 20 Cases of Pre-Eclampsia. Softpedia; 2013 Feb 8. Available from: https:&#47;&#47;news.softpedia.com&#47;news&#47;Air-Pollution-Linked-to-1-in-20-Cases-of-Pre-Eclampsia-328058.shtml</RefTotal>
        <RefLink>https:&#47;&#47;news.softpedia.com&#47;news&#47;Air-Pollution-Linked-to-1-in-20-Cases-of-Pre-Eclampsia-328058.shtml</RefLink>
      </Reference>
      <Reference refNo="18">
        <RefAuthor>Sj&#248;gren K</RefAuthor>
        <RefTitle>Traffic noise and pollution increase risk of pre-eclampsia during pregnancy</RefTitle>
        <RefYear>2016</RefYear>
        <RefJournal>Science Nordic</RefJournal>
        <RefPage></RefPage>
        <RefTotal>Sj&#248;gren K. Traffic noise and pollution increase risk of pre-eclampsia during pregnancy. Science Nordic; 2016 Oct 27. Available from: https:&#47;&#47;www.sciencenordic.com&#47;denmark-environment-videnskabdk&#47;traffic-noise-and-pollution-increase-risk-of-pre-eclampsia-during-pregnancy&#47;1438950</RefTotal>
        <RefLink>https:&#47;&#47;www.sciencenordic.com&#47;denmark-environment-videnskabdk&#47;traffic-noise-and-pollution-increase-risk-of-pre-eclampsia-during-pregnancy&#47;1438950</RefLink>
      </Reference>
      <Reference refNo="19">
        <RefAuthor>Galvin G</RefAuthor>
        <RefTitle>Air pollution tied to hypertension in pregnant women</RefTitle>
        <RefYear>2019</RefYear>
        <RefJournal>US News &#38; World Report</RefJournal>
        <RefPage></RefPage>
        <RefTotal>Galvin G. Air pollution tied to hypertension in pregnant women. US News &#38; World Report; 2019 Dec 18. Available from: 
https:&#47;&#47;www.usnews.com&#47;news&#47;healthiest-communities&#47;articles&#47;2019-12-18&#47;air-pollution-tied-to-hypertension-in-pregnant-women-study</RefTotal>
        <RefLink>https:&#47;&#47;www.usnews.com&#47;news&#47;healthiest-communities&#47;articles&#47;2019-12-18&#47;air-pollution-tied-to-hypertension-in-pregnant-women-study</RefLink>
      </Reference>
      <Reference refNo="20">
        <RefAuthor>McHugh E</RefAuthor>
        <RefTitle>Clean the Air: Protecting Expectant Mothers from the Hazards of Pollution</RefTitle>
        <RefYear>2024</RefYear>
        <RefJournal>MedpageToday</RefJournal>
        <RefPage></RefPage>
        <RefTotal>McHugh E. Clean the Air: Protecting Expectant Mothers from the Hazards of Pollution. MedpageToday; 2024 Jun 15. Available from: https:&#47;&#47;www.medpagetoday.com&#47;opinion&#47;second-opinions&#47;110667</RefTotal>
        <RefLink>https:&#47;&#47;www.medpagetoday.com&#47;opinion&#47;second-opinions&#47;110667</RefLink>
      </Reference>
      <Reference refNo="21">
        <RefAuthor>HealthDay staff</RefAuthor>
        <RefTitle>Pregnant Woman Exposed to 45 Common Chemicals. Study Finds</RefTitle>
        <RefYear>2026</RefYear>
        <RefJournal>US News &#38; World Reports, HealthDay News</RefJournal>
        <RefPage></RefPage>
        <RefTotal>HealthDay staff. Pregnant Woman Exposed to 45 Common Chemicals. Study Finds. US News &#38; World Reports, HealthDay News; 2026 Jun 17. Available from: https:&#47;&#47;www.usnews.com&#47;news&#47;health-news&#47;articles&#47;2026-06-17&#47;pregnant-woman-exposed-to-45-common-chemicals-study-finds</RefTotal>
        <RefLink>https:&#47;&#47;www.usnews.com&#47;news&#47;health-news&#47;articles&#47;2026-06-17&#47;pregnant-woman-exposed-to-45-common-chemicals-study-finds</RefLink>
      </Reference>
      <Reference refNo="22">
        <RefAuthor>Epidemiology</RefAuthor>
        <RefTitle>Impact of Road Traffic Pollution on Pre-eclampsia and Pregnancy-induced Hypertensive Disorders &#91;Video&#93;</RefTitle>
        <RefYear>2017</RefYear>
        <RefTotal>Epidemiology. Impact of Road Traffic Pollution on Pre-eclampsia and Pregnancy-induced Hypertensive Disorders &#91;Video&#93;. YouTube; 2017. Available from: https:&#47;&#47;www.youtube.com&#47;watch&#63;v&#61;C6moDXFI47U</RefTotal>
        <RefLink>https:&#47;&#47;www.youtube.com&#47;watch&#63;v&#61;C6moDXFI47U</RefLink>
      </Reference>
      <Reference refNo="23">
        <RefAuthor>FIU in DC</RefAuthor>
        <RefTitle>Air pollution and Pre-eclampsia &#91;Video&#93;</RefTitle>
        <RefYear>2022</RefYear>
        <RefTotal>FIU in DC. Air pollution and Pre-eclampsia &#91;Video&#93;. YouTube; 2022 Aug 28. Available from: https:&#47;&#47;www.youtube.com&#47;watch&#63;v&#61;aN8NX3ZUiPM</RefTotal>
        <RefLink>https:&#47;&#47;www.youtube.com&#47;watch&#63;v&#61;aN8NX3ZUiPM</RefLink>
      </Reference>
      <Reference refNo="24">
        <RefAuthor>Kumar A</RefAuthor>
        <RefTitle>How does pollution affect pregnancy &#91;Video&#93;</RefTitle>
        <RefYear>2025</RefYear>
        <RefTotal>Kumar A. How does pollution affect pregnancy &#91;Video&#93;. YouTube; 2025 Nov 28. Available from: https:&#47;&#47;www.youtube.com&#47;shorts&#47;WI85EZil26Y</RefTotal>
        <RefLink>https:&#47;&#47;www.youtube.com&#47;shorts&#47;WI85EZil26Y</RefLink>
      </Reference>
      <Reference refNo="25">
        <RefAuthor>CBS New York</RefAuthor>
        <RefTitle>Max Minute: New Data Reveals Damaging Effects Of Climate Change On Pregnant Women &#91;Video&#93;</RefTitle>
        <RefYear>2020</RefYear>
        <RefTotal>CBS New York. Max Minute: New Data Reveals Damaging Effects Of Climate Change On Pregnant Women &#91;Video&#93;. YouTube; 2020 Jun 18. Available from: https:&#47;&#47;www.youtube.com&#47;watch&#63;v&#61;cQFGqLeaGug</RefTotal>
      </Reference>
      <Reference refNo="26">
        <RefAuthor>News9Live</RefAuthor>
        <RefTitle>Invisible Danger: Air Pollution&#8217;s Impact on Pregnancy &#91;Video&#93;</RefTitle>
        <RefYear>2025</RefYear>
        <RefTotal>News9Live. Invisible Danger: Air Pollution&#8217;s Impact on Pregnancy &#91;Video&#93;. YouTube; 2025 Dec 18. Available from: https:&#47;&#47;www.youtube.com&#47;watch&#63;v&#61;8k00wKh4Drg</RefTotal>
      </Reference>
      <Reference refNo="27">
        <RefAuthor>Salazar M</RefAuthor>
        <RefAuthor>Chaw M</RefAuthor>
        <RefAuthor>Hellier Y</RefAuthor>
        <RefAuthor>Hsia S</RefAuthor>
        <RefAuthor>Gruenberg K</RefAuthor>
        <RefTitle>Comparison of Qualitative Analyses Conducted by Artificial Intelligence Versus Traditional Methods</RefTitle>
        <RefYear>2025</RefYear>
        <RefJournal>Am J Pharm Educ</RefJournal>
        <RefPage>101882</RefPage>
        <RefTotal>Salazar M, Chaw M, Hellier Y, Hsia S, Gruenberg K. Comparison of Qualitative Analyses Conducted by Artificial Intelligence Versus Traditional Methods. Am J Pharm Educ. 2025 Dec;89(12):101882. DOI: 10.1016&#47;j.ajpe.2025.101882</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1016&#47;j.ajpe.2025.101882</RefLink>
      </Reference>
      <Reference refNo="28">
        <RefAuthor>Lian X</RefAuthor>
        <RefAuthor>Krichene W</RefAuthor>
        <RefAuthor>Huang B</RefAuthor>
        <RefAuthor></RefAuthor>
        <RefTitle>Quantization inflates reasoning: token inflation as a hidden cost of low-bit reasoning models</RefTitle>
        <RefYear>2026</RefYear>
        <RefJournal>arXiv</RefJournal>
        <RefPage></RefPage>
        <RefTotal>Lian X, Krichene W, Huang B, et al. Quantization inflates reasoning: token inflation as a hidden cost of low-bit reasoning models. arXiv. 2026. DOI: 10.48550&#47;arXiv.2606.25519</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.48550&#47;arXiv.2606.25519</RefLink>
      </Reference>
      <Reference refNo="29">
        <RefAuthor>Chochlakis G</RefAuthor>
        <RefAuthor>Pandiyan NM</RefAuthor>
        <RefAuthor>Lerman K</RefAuthor>
        <RefAuthor>Narayanan S</RefAuthor>
        <RefTitle>Larger language models don&#8217;t care how you think: why chain-of-thought prompting fails in subjective tasks</RefTitle>
        <RefYear>2024</RefYear>
        <RefBookTitle>IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2025 Apr 6-11; Hyderabad, India</RefBookTitle>
        <RefPage></RefPage>
        <RefTotal>Chochlakis G, Pandiyan NM, Lerman K, Narayanan S. Larger language models don&#8217;t care how you think: why chain-of-thought prompting fails in subjective tasks. In: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2025 Apr 6-11; Hyderabad, India. arXiv; 2024. 
DOI: 10.48550&#47;arXiv.2409.06173</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.48550&#47;arXiv.2409.06173</RefLink>
      </Reference>
      <Reference refNo="30">
        <RefAuthor>Hitch D</RefAuthor>
        <RefTitle>Artificial intelligence augmented qualitative analysis: The way of the future&#63; Qualitative Health Research 2024</RefTitle>
        <RefYear></RefYear>
        <RefTotal>Hitch D. Artificial intelligence augmented qualitative analysis: The way of the future&#63; Qualitative Health Research 2024;34(7):595-606. DOI: 10.1177&#47;10497323231217392</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1177&#47;10497323231217392</RefLink>
      </Reference>
      <Reference refNo="31">
        <RefAuthor>Sen M</RefAuthor>
        <RefAuthor>Sen SN</RefAuthor>
        <RefAuthor>Sahin TG</RefAuthor>
        <RefTitle>A new era for data analysis in qualitative research: ChatGPT&#33; Shanlax International Journal of Education</RefTitle>
        <RefYear>2023</RefYear>
        <RefTotal>Sen M, Sen SN, Sahin TG. A new era for data analysis in qualitative research: ChatGPT&#33; Shanlax International Journal of Education. 2023;11:1-15.</RefTotal>
      </Reference>
      <Reference refNo="32">
        <RefAuthor>Nguyen DC</RefAuthor>
        <RefAuthor>Welch C</RefAuthor>
        <RefTitle>Generative artificial intelligence in qualitative data analysis: analyzing -or just chatting&#63; Organizational Research Methods</RefTitle>
        <RefYear>2026</RefYear>
        <RefTotal>Nguyen DC, Welch C. Generative artificial intelligence in qualitative data analysis: analyzing -or just chatting&#63; Organizational Research Methods. 2026;29(1):3-39. 
DOI: 10.1177&#47;10944281251377154</RefTotal>
        <RefLink>https:&#47;&#47;doi.org&#47;10.1177&#47;10944281251377154</RefLink>
      </Reference>
    </References>
    <Media>
      <Tables>
        <Table format="png">
          <MediaNo>1</MediaNo>
          <MediaID>1</MediaID>
          <Caption><Pgraph><Mark1>Table 1: Prompts to perform coding with ChatGPT 4.1 without internet access</Mark1></Pgraph></Caption>
        </Table>
        <Table format="png">
          <MediaNo>2</MediaNo>
          <MediaID>2</MediaID>
          <Caption><Pgraph><Mark1>Table 2: Coding based on complete abstracts and exemplary words found in at least nine abstracts and additional prevalent words appearing in fewer abstracts</Mark1></Pgraph></Caption>
        </Table>
        <Table format="png">
          <MediaNo>3</MediaNo>
          <MediaID>3</MediaID>
          <Caption><Pgraph><Mark1>Table 3: Codes generated by direct thematic coding from AI (Strategy 8a) compared to those generated by manual coding of prevalent words</Mark1></Pgraph></Caption>
        </Table>
        <Table format="png">
          <MediaNo>4</MediaNo>
          <MediaID>4</MediaID>
          <Caption><Pgraph><Mark1>Table 4: Summary of thematic coding using Strategy 8b</Mark1></Pgraph></Caption>
        </Table>
        <NoOfTables>4</NoOfTables>
      </Tables>
      <Figures>
        <Figure width="427" height="376" format="png">
          <MediaNo>1</MediaNo>
          <MediaID>1</MediaID>
          <Caption><Pgraph><Mark1>Figure 1: ROC curve for prompting strategies applied to structured abstracts</Mark1></Pgraph></Caption>
        </Figure>
        <Figure width="428" height="375" format="png">
          <MediaNo>2</MediaNo>
          <MediaID>2</MediaID>
          <Caption><Pgraph><Mark1>Figure 2: ROC for prompting strategies applied to narratives</Mark1></Pgraph></Caption>
        </Figure>
        <NoOfPictures>2</NoOfPictures>
      </Figures>
      <InlineFigures>
        <NoOfPictures>0</NoOfPictures>
      </InlineFigures>
      <Attachments>
        <NoOfAttachments>0</NoOfAttachments>
      </Attachments>
    </Media>
  </OrigData>
</GmsArticle>