Toggle navigation
Home
▼ Details
Products and pricing
Chart gallery
User stories
Text analytics
CDC NAMCS Library
Blog
Tutorials
Contact
Sign in
Post Editor
← All help posts
View post
Save
*This article walks through how to evaluate Open-end coding quality in Protobi before delivering results.* [Image: Showing the purpose of this article] ## Coding Quality Open-ended survey questions allow respondents to share detailed, unscripted feedback in their own words. Analyzing these responses requires coding — assigning each response to one or more thematic categories so they can be counted and compared across the dataset. Protobi supports coding both manually and through Autocode, which uses AI to automatically assign responses to categories — and two questions determine whether those results are ready to deliver: - **Are the categories the right ones for the question being asked?** - **Were individual responses assigned to the correct categories?** <img src="/uploads/images/image-evaluate-open-end-coding-quality-2026-06-01-20-18-08.png" alt="showing responses on the left and assigned codes on the right" style="max-width: 700px"/> A few dimensions matter when assessing coding quality: - **Category validity (i.e code)** — each category should represent a real, recurring theme, distinct from the others, with a name and definition that matches what the assigned responses actually say - **Assignment accuracy** — each response lands in the right category, with the same logic applied consistently across the full dataset. - **Inter-coder reliability** — the degree of agreement among independent evaluators when they code, rate, or assess the same data. In open-ended survey coding, high inter-coder reliability means multiple coders applying the same scheme reach the same conclusions independently. ## Measuring Coding Concordance Protobi checks all these dimensions by comparing two independent coding solutions: - Which assignments appear in both solutions vs. just one - Which assignments have higher vs. lower concordance - Whether individual code definitions and assignments hold up on their own terms ### Comparing Two Coding Solutions When two coding solutions are applied to the same set of responses, some responses will land in the same category in both. Others will appear in one solution but not the other. Comparing the two reveals where the solutions agree, where they diverge, and why. Traditionally, overlap between two solutions was shown using pie charts — useful for seeing 2-3 segments but limited when comparing across an entire codeset. Protobi uses a Sankey diagram instead — a Mastercard-like structure where each code appears on both sides and the overlap in the middle shows how many responses both solutions agreed on, across all codes at once. #### Same Codesets When both solutions use the same codeset, the comparison is direct. Take "Ease of Use" — how many responses coded as "Ease of Use" in Solution 1 are also coded as "Ease of Use" in Solution 2? Protobi displays this as a Sankey diagram — each code appears on both sides, and the overlap in the middle shows how many responses both solutions assigned to the same code. The wider the overlap, the stronger the concordance. Three patterns are worth looking for: **High overlap.** The same responses appear under this code in both solutions. The code is being applied consistently. **Different volumes, same name.** One solution assigned many more responses to this code than the other. Both solutions have the code but draw its boundaries differently — one applies it broadly, the other narrowly. **Similar volumes, but different responses.** Both solutions have roughly the same count under this code, but not the same responses. The two solutions disagree on which responses actually belong there. The code definition likely needs tightening. <img src="/uploads/images/image-evaluate-open-end-coding-quality-2026-06-02-19-32-25.png" alt="Sankey diagram showing two solutions with the same codeset — wide overlap on Ease of Use, asymmetric lengths on others" style="max-width: 650px"/> #### Different Codesets When coding solutions are developed independently, different names often emerge for the same underlying theme. One analyst might call a code "Complexity/Overwhelming," another might call it "Complexity Concerns." The Sankey diagram shows whether responses grouped under one name in Solution 1 are the same responses grouped under a different name in Solution 2 — making it possible to tell whether the difference is labeling or a genuine coding disagreement. Four things to look for: - **Same responses, different names** — both solutions captured the same idea. The difference is labeling, not substance. - **Different names, very different response counts** — one solution applied this theme more broadly than the other, even after accounting for the name difference. - **A code in one solution with no equivalent in the other** — one solution identified a theme the other missed entirely. - **One code in one solution mapping to multiple in the other** — one solution merged themes the other kept separate. The structural difference needs to be understood before any downstream metrics are compared. <img src="/uploads/images/image-evaluate-open-end-coding-quality-2026-06-02-19-31-48.png" alt="Sankey diagram showing two solutions with different codesets — responses flowing between differently named but conceptually related codes" style="max-width: 780px"/> ### Assignment-Level Comparison **In both current and reference.** > *"Looks easy to use. Practices are driven by price. I am eco friendly so I like reusable."* > > Both solutions assign: Easy to use, Cost effective / pricing concerns, Environmentally friendly / reusable **In reference only (+Right).** > *"Looks good but would have to be cheaper than Product B."* > > Current coding assigns: Generally positive impression, Cost effective / pricing concerns > > Reference assigns: Generally positive impression, Cost effective / pricing concerns, Comparable to existing products > > *Comparable to existing products* is missing from the current coding. **In current coding only (+Left).** > *"Looks useful and superior to current devices — like that the medicine clears."* > > Current coding assigns: Innovative / improvement over alternatives, Other > > Reference assigns: Innovative / improvement over alternatives > > *Other* was not assigned by the reference. *[Image: Response-level comparison view showing a response with Match, +Right, and +Left indicators side by side]* Codes with consistently high +Right point to themes the current coding missed. Codes with consistently high +Left point to codes the current coding over-applied. ## Independent Coding Review ### Reviewing Code Definitions **Does the label match the definition?** If the code label and its definition do not match, coders will assign it to the wrong responses. **Is it complete?** A definition covering only a narrow subset leaves edge cases open — different coders will interpret the same response differently. **Is it specific enough?** A definition too broad leads different coders to draw its boundaries differently. **Does it match the survey question?** A code capturing a recurring theme unrelated to what was asked may reflect respondent tangents rather than meaningful signal. *[Image: Code set panel in Recode (Advanced) showing code labels and definitions]* ### Reviewing Value-Code Assignments - Does this response actually say what the code claims it says? - If the response covers multiple themes, did the coding capture all of them — and does any code cover something only faintly implied rather than clearly stated? In Protobi, use the Validate workflow to work through assignments response by response and record a decision on each one before finalizing. ## Resolving Match and Mismatch Assignments Once the comparison is complete, each discrepancy needs a decision before the project is ready to deliver. In Protobi, use the QC column to record a decision for each discrepancy: - **Yes** — the assignment is correct. Keep it as is. - **No** — the assignment is incorrect. Unapply or reassign the response to a more accurate code. - **TBD** — needs a closer look. Resolve it as Yes or No before delivery. *[Image: QC column showing Yes, No, and TBD decisions next to flagged assignments in the comparison view]* ## Evaluating Coding Quality in Protobi Protobi provides two workflows for evaluating coding quality. **Evaluate** generates an independent reference solution by re-applying the existing coding scheme to all responses without referencing the existing assignments. It then compares the two solutions response by response and identifies differences. Protobi also automatically reviews each discrepancy — assessing whether the assignment is reasonable given the raw response text and the code definition. **Validate** supports manual spot-checking. Go through each response and its assigned code, decide whether the assignment is correct, and record a Yes, No, or TBD decision in the QC column for each one. ### Evaluate Assignments **1. Start the Evaluate process** Evaluate compares two coding solutions. These can be: - Two human-coded solutions - One human-coded and one AI-generated solution - Two AI-generated solutions — for example, generate one solution using Autocode first, then run Evaluate to generate a second independent solution for comparison If a second solution does not exist yet, Evaluate will generate one by re-applying the existing coding scheme independently. Access Evaluate in one of two ways: - Click the **bot icon** on any coded element — the Autocode AI dialog will appear with the **Evaluate** option - In **Recode (Advanced)**, select **Autocode with Protobi AI** from the lower right corner, then choose **Evaluate** Click **OK** to begin. **2. Analysis and comparison** Protobi compares the two coding solutions — the existing assignments against the reference solution — and identifies every discrepancy between them. **3. Review assignments using the comparison toggle indicators** - **In both schemes** — both solutions agree on this assignment. No action needed. - **In current scheme only** — this assignment exists in the current coding but not in the reference. Review and decide whether to keep or remove it. - **In reference scheme only** — this assignment exists in the reference but not in the current coding. Review and decide whether to add it. *[Image: Evaluate comparison view showing toggle indicators and QC column]* **4. Record your QC decisions** Record a decision for each discrepancy in the QC column — **Yes, No, or TBD**. **5. Make adjustments and save** - Use the action buttons in the comparison panel to accept or reject each discrepancy. - Reassign all **No** responses to a more accurate code. - Resolve all **TBD** assignments before finalizing. - Press **Save** in the project toolbar to save all changes. *[Image: Comparison panel showing action buttons for accepting and rejecting assignments]* ### Validate Existing Assignments > **Note:** Complete coding first — using Autocode or manually — then use Validate to spot-check assignments. **1. Review each assignment** - Go through each response and its assigned code in Recode (Advanced). - Ask: does this code accurately represent what the respondent said? **2. Record your QC decisions** Record a decision for each assignment in the QC column — **Yes, No, or TBD** **3. Make adjustments and save** - Unapply or reassign all **No** assignments. - Resolve all **TBD** assignments before finalizing. - Press **Save** in the project toolbar to save all changes. *[Image: QC column showing Yes/No/TBD values next to coded responses in Recode (Advanced)]* ## Background Research Two broad approaches exist in academic literature for measuring coding quality, depending on whether categories already exist or need to be developed from the data. **When categories already exist (deductive coding)**, quantitative metrics measure agreement between two independent solutions: **Observed agreement** — the percent of responses receiving the same code in both solutions. Scale 0% to 100%. **Cohen's Kappa** — observed agreement scaled relative to the likelihood of agreement by chance. Works with exactly two raters and exclusive categories. Scale -1 to +1. **Krippendorff's Alpha** — adjusts observed agreement relative to the likelihood of agreement by chance. Works with multiple raters and overlapping categories. Scale 0 to +1. **Gwet's AC1** — resolves paradoxes in Krippendorff's Alpha by modeling chance agreement versus agreement on obvious cases. Works with multiple raters and overlapping categories. Scale 0 to +1. **When analysts develop categories from the data (inductive coding)**, qualitative trustworthiness criteria apply instead of numeric metrics. These include credibility, transferability, dependability, and confirmability — but do not produce quantitative measures of the coding scheme itself. Current methods have a shared limitation: they either provide a quantitative metric but assume a gold standard solution already exists to compare against, or focus on how categories are developed without measuring the coding scheme quantitatively. Protobi's framework addresses both. It generates its own independent reference solution — so analysts do not need a pre-existing gold standard — and measures agreement at the assignment level across multi-label coding schemes, applying Kappa per code-response pair rather than across the full dataset. In a pharma study, this approach returned 83.02% assignment agreement and 84.8% response-level agreement across 99 responses. ### Related guides: [Autocode with Protobi AI →](https://help.protobi.com/content/help/new-to-protobi/autocode-with-protobi-ai_2) [Recode (Advanced) guide →](https://help.protobi.com/content/help/text-open-end-questions/reformat-tool-advanced) [Simple Recode guide →](https://help.protobi.com/content/help/new-to-protobi/simple-open-end-recoding) [Coding Open-end Responses in Protobi →](https://help.protobi.com/content/help/new-to-protobi/coding-open-end-responses-in-protobi)
Publishing
Date
Status
Published
Draft
Slug
edit
Content
Thumbnail
Categories
Manage
New to Protobi?
Charts
Making Changes
Intermediate topics for editors
Frequently Asked Questions
SERMO Topics
Tutorial Pages
Internal Docs
Data Processing
Videos
Obsolete
GSG Admin
SERMO Admin
GSG Topics
How to...
Basics for viewers
Basics for editors
For project admins
Advanced topics
Assessments
Articles (in-progress)
New and updated articles
Process data in Protobi
superseded
Text open-end questions
Troubleshooting
Tracking studies
Organizing the view
API References
Tools
AI Database
Checking AI...
Convert to MD
Danger zone
Delete