End-to-End Text-Analysis Project

Weight 50% of final grade
Components Topic brief, poster, and GitHub repository
Work Format 1–4 people
Submission Canvas

Complete an end-to-end text-analysis project using a real-world collection such as news articles, social media posts, policy documents, or another appropriate source. The project should transform raw, unstructured text into analysis-ready data; implement a reproducible Python workflow for a course-aligned task; assess the quality and limitations of the results; and communicate the findings clearly through a poster and brief verbal explanation. Appropriate applications include classification, clustering, topic modeling, semantic search, retrieval-augmented generation, and question answering.

You may complete the project individually or in a group of up to four people. Group work is highly encouraged.

Begin by submitting a Topic Brief in one shared project repository. Continue using that same repository as you develop the full project. Use Python and, where appropriate, Google Colab and Hugging Face. Submit the Topic Brief repository URL, poster, and final GitHub repository URL through their designated Canvas links; work submitted by email will not be graded. The poster is due Wednesday, November 25, 2026, at 11:59 p.m. Late work is accepted with a 25% deduction for each day submitted late.

Component Weights

Project components and their percentages of the final course grade
Component Percent of Final Grade
Topic Brief 5%
Project poster presentation 25%
Documented GitHub repository 20%
Component 1

Topic Brief

Due date: Saturday, October 17, 2026, at 11:59 p.m.

Work with your project team to create one shared GitHub repository and use its root README.md to introduce your proposed text-analysis project. The brief should make the task, intended output, preferred data, and immediate next steps understandable enough for the instructor to give useful feedback. This is an initial plan: you may revise the topic, data, or methods as the project develops.

Required README sections

  1. Working project title: Use a specific title that distinguishes the project. The title may change as the project develops.
  2. Team members: List the full name of every person working in the shared repository.
  3. One-sentence task: State one specific operation in the form “Given [input], we will [analyze, classify, retrieve, compare, or generate] [output].” Describe what the project will do, not only its general topic.
  4. Illustrative input and desired output: Provide one representative, non-sensitive input and show or describe what the proposed system should return, predict, rank, extract, compare, or generate.
  5. Background and starting point: Briefly explain why the team is interested in the topic and which examples, experiences, questions, applications, or resources led to the idea. A literature review or extended theoretical framing is not required.
  6. Intended output and success: Describe what the project should produce or accomplish, who would use the output when relevant, and—in one sentence—what would make the output useful or successful. A formal metric is not required at this stage.
  7. Preferred data: Identify the proposed corpus, its original source and collection process, direct link, access method, unit of analysis, relevant text or fields, approximate size, and any access, licensing, privacy, or redistribution concerns. Include or link to two or three representative, non-sensitive examples, explain how the corpus supports the task, and identify an optional backup corpus when access is uncertain.
  8. Feasibility and immediate next steps: Define the minimum viable project, report whether a team member has accessed the data, identify the largest current risk and a backup plan, and assign the next two actions to team members with approximate target dates.
  9. Possible method and evaluation (optional): If the team already has ideas, identify possible text-analysis methods and how you might determine whether the output performs well or is useful. It is acceptable to leave this undecided.
  10. Questions for the instructor: Ask for the feedback your team needs about scope, data, methods, evaluation, tools, or risks. Write “No additional questions at this time” if none remain.
  11. AI-use statement: Identify any generative or agentic AI tool used to prepare the brief, how it was used, what work preceded it, which advice was incorporated, and how the result was verified. State clearly if no such tool was used.

One repository per team: Download the Markdown template, place it at the root of the team’s project repository as README.md, and submit that repository’s URL through Canvas. Continue using the same repository for the final project. The repository must be accessible to the instructor.

Zero-credit naming rule: The repository name must begin with the case-sensitive prefix DSA495, such as DSA495-news-classification. Check the name before submitting the URL.

How credit is earned

The Topic Brief is graded for completion and specificity. Full credit requires the correct repository name, one accessible team repository with the completed template at its root, a concrete task with an illustrative input and desired output, a clear success statement, a sufficiently detailed and accessible preferred corpus, a feasible minimum project and immediate action plan, thoughtful questions or an explicit statement that the team has none, and a complete AI-use statement. The optional methods and evaluation section is not required for full credit.

Evaluation

How the Topic Brief is evaluated

These criteria total 100% of the Topic Brief component.

30%

Task Definition & Intended Output

Specificity of the input, analytical operation, output, illustrative example, intended use, and initial definition of success.

35%

Corpus Specification & Traceability

Clarity of the corpus, original source and collection, verified access, unit of analysis, relevant fields, examples, scale, and data-use constraints.

20%

Scope & Feasibility

Fit between the task and data, a realistic minimum project, recognition of consequential risks, and concrete assigned next actions.

15%

Repository Readiness & Completeness

Correct repository setup, required README sections, team identification, accessibility, source links, and AI-use documentation.

View the Topic Brief rubric Expand to compare performance descriptions for all four criteria.

The performance band indicates the percentage of points earned within a criterion; the Weight column determines that criterion’s contribution to the component grade.

Detailed evaluation rubric for the Topic Brief component
Criterion Exemplary
95–100%
Proficient
80–94.9%
Developing
40–79.9%
Needs Improvement
Below 40%
Weight
Task Definition & Intended Output The one-sentence task identifies a concrete input, operation, and output. The illustrative example makes the proposed behavior clear, and the intended result, use, and initial definition of success are precise, coherent, and technically plausible. The task, illustrative example, intended output, and success statement are clear overall. Minor ambiguity remains about the input, operation, output format, intended use, or what would count as success. The general topic is apparent, but the task is broad or the illustrative example, intended output, intended use, or success statement is incomplete or unclear. The submission names a topic without defining an actionable text-analysis task or meaningful intended output. 30%
Corpus Specification & Traceability The corpus is identified with its original source and collection process, a direct link, a documented access check, and representative examples. Unit of analysis, relevant text and fields, approximate scale, task fit, and licensing or privacy constraints are documented accurately. The corpus, provenance, access status, and relationship to the task are suitable and understandable, with only minor gaps in examples, collection details, fields, scale, or usage constraints. A possible dataset is named, but its original source, collection, contents, access status, examples, or relationship to the task remains insufficiently verified. No usable corpus is identified, the source cannot be verified, or the proposed data does not contain the information required by the task. 35%
Scope & Feasibility The task and corpus fit one another and support a realistic minimum viable project. The brief identifies the largest consequential risk, provides a workable backup plan, and assigns two concrete next actions with owners and approximate target dates. The proposal appears feasible and appropriately scoped overall, with a plausible minimum project and action plan. A risk, backup plan, owner, or target date could be specified more clearly. The idea may be workable, but the minimum project, major risk, backup plan, or immediate responsibilities are incomplete, leaving unresolved scope, access, labeling, computing, or evaluation concerns. The proposed work is not feasible with the identified data, time, or course tools, and the brief does not identify a viable minimum version or executable next step. 20%
Repository Readiness & Completeness The accessible shared repository follows the required naming rule, places a complete README at the root, identifies all team members, links sources, and includes a complete AI-use statement. The repository and README meet the core requirements, with only minor omissions or organization issues that do not impede review. The repository is accessible, but several required sections, links, team details, or documentation elements are incomplete or difficult to locate. The repository is inaccessible or substantially incomplete. A repository name that does not begin with the case-sensitive prefix DSA495 receives zero credit for the Topic Brief. 15%
Component 2

Project Poster Presentation

Poster due: Wednesday, November 25, 2026, at 11:59 p.m. · Present in class: Monday, November 30, 2026

Present the project in a poster format as a concise, audience-centered explanation of the analytical purpose, text data, technical workflow, principal findings, and practical implications. Use readable and accurate visualizations to make the text-analysis results accessible without obscuring important limitations.

Suggested poster structure

  1. Frame the projectIntroduce the text-analysis question or application and explain why it matters.
  2. Explain the workflowDescribe the text collection, preparation decisions, representation, and modeling approach.
  3. Show the resultsUse clear visuals or examples to communicate the most important findings and model behavior.
  4. Evaluate and reflectDiscuss quality, limitations, responsible use, and implications for interpretation or decision-making.

Be prepared to explain the code, modeling decisions, evaluation strategy, and limitations of the project and to answer questions about how the submitted work was produced. Include the required AI-use statement in the submitted materials.

Evaluation

How the Poster Presentation is evaluated

These criteria total 100% of the Poster Presentation component.

25%

Problem & Technical Workflow

Precision of the task, data pipeline, representations, baselines, models, and important implementation decisions.

30%

Experimental Evidence & Evaluation Validity

Quality of the comparison design, metrics, test evidence, error analysis, and attention to leakage, uncertainty, and limitations.

25%

Poster Design & Technical Communication

Accuracy, readability, visual hierarchy, purposeful figures, and the ability to communicate technical evidence concisely.

20%

Oral Defense & Team Command

Ability to explain design decisions, code, results, failure modes, and individual or team contributions and to answer questions accurately.

View the Poster Presentation rubric Expand to compare performance descriptions for all four criteria.

The performance band indicates the percentage of points earned within a criterion; the Weight column determines that criterion’s contribution to the component grade.

Detailed evaluation rubric for the Poster Presentation component
Criterion Exemplary
95–100%
Proficient
80–94.9%
Developing
40–79.9%
Needs Improvement
Below 40%
Weight
Problem & Technical Workflow The poster defines the computational task precisely and presents a coherent end-to-end workflow. Data preparation, representations, baselines, models, and consequential implementation choices are technically sound and easy to trace. The task and workflow are correct and understandable overall, with minor omissions in implementation detail, rationale, or connections between stages. The main approach is visible, but important pipeline stages, baselines, model choices, or implementation decisions are incomplete, unclear, or technically questionable. The task or workflow is substantially unclear or incorrect, and the audience cannot determine how the system produced its results. 25%
Experimental Evidence & Evaluation Validity The evaluation design matches the task and uses appropriate baselines, splits, metrics, examples, and error analysis. Claims follow from the evidence, and leakage, uncertainty, failure modes, and important limitations are handled rigorously. Evaluation evidence is appropriate and supports the main claims, with only minor gaps in comparisons, metrics, error analysis, or discussion of validity threats. Some results are reported, but weak baselines, limited test evidence, inappropriate metrics, possible leakage, or incomplete error analysis reduces confidence in the conclusions. Evaluation is absent, invalid, or misleading, or claims about system quality are unsupported by reproducible evidence. 30%
Poster Design & Technical Communication The poster has a strong visual hierarchy and uses concise text, readable figures, accurate labels, and well-selected examples to communicate the technical work and results without oversimplification. The poster is clear and readable overall. Visuals and explanations communicate the main technical story, with minor issues in density, labeling, organization, or emphasis. The main message can be found, but crowded text, weak organization, default or unclear visuals, missing labels, or imprecise explanations make the technical evidence difficult to follow. The poster is incomplete, inaccurate, or visually inaccessible and does not communicate the project’s task, workflow, evidence, and conclusions effectively. 25%
Oral Defense & Team Command The presenter or team explains design decisions, implementation, results, and limitations accurately and concisely. Responses to questions demonstrate direct command of the submitted code and clear individual contributions. The presenter or team explains the project and answers questions accurately overall, with minor uncertainty about details or contributions. Explanations show partial understanding, but important design decisions, results, code behavior, or team contributions cannot be explained clearly. The presenter or team cannot explain essential parts of the submitted work, answer basic technical questions, or establish meaningful participation in the project. 20%
Component 3

Documented GitHub Repository

Due date: December 6, 2026

Submit a well-organized GitHub repository that contains the complete project workflow, from raw text through interpretable results. The repository may include multiple notebooks, scripts, configuration files, and supporting materials, but it must provide a clear entry point and enough documentation for another reader to understand and reproduce the work.

Required elements

  • README and project map: State the project purpose, identify the recommended starting point, explain the repository structure, and provide step-by-step instructions for reproducing the analysis.
  • Environment and data access: Document required packages, software versions, setup steps, data sources, and any instructions needed to obtain or reconstruct the input data without committing restricted material.
  • Preparation and management: Show how raw text was collected, cleaned, organized, tokenized, and otherwise prepared for analysis.
  • Modeling or application: Implement a course-aligned approach such as classification, clustering, topic modeling, semantic search, retrieval-augmented generation, or question answering.
  • Evaluation and interpretation: Assess output quality, document important errors or failure modes, explain limitations, and include clear visual evidence supporting the conclusions.
  • Repository quality: Use meaningful file names and folders, relative paths, concise comments, and version-controlled source files so the complete workflow can be rerun from beginning to end.

Include an AI-use statement in the repository. Follow the course AI policy, retain evidence of your process, protect confidential or restricted data, and be prepared to explain and reproduce your code, analytical choices, results, and any AI-assisted work.

Evaluation

How the GitHub Repository is evaluated

These criteria total 100% of the Documented GitHub Repository component.

30%

Software Correctness & Complete Workflow

Correct implementation of the complete pipeline from source data through preparation, modeling or application, and final output.

25%

Reproducibility & Environment Management

Reliable setup instructions, pinned or documented dependencies, portable paths, data access procedures, and deterministic execution where appropriate.

25%

Evaluation & Testing Evidence

Task-appropriate baselines, metrics, validation logic, saved outputs, error analysis, and checks that support claims about system behavior.

20%

Code Quality, Documentation & Responsible Engineering

Readable modular code, clear repository structure, useful documentation, data and model provenance, limitations, and transparent AI use.

View the GitHub Repository rubric Expand to compare performance descriptions for all four criteria.

The performance band indicates the percentage of points earned within a criterion; the Weight column determines that criterion’s contribution to the component grade.

Detailed evaluation rubric for the Documented GitHub Repository component
Criterion Exemplary
95–100%
Proficient
80–94.9%
Developing
40–79.9%
Needs Improvement
Below 40%
Weight
Software Correctness & Complete Workflow The repository implements a complete, technically sound pipeline from source data through final output. Core code executes successfully, algorithms are used correctly, and outputs are consistent with the stated task. The complete workflow is present and correct overall. Minor implementation issues or omissions do not materially change the main results. Major stages are present, but incomplete code, technical errors, hard-coded assumptions, or disconnected steps limit the workflow’s correctness or completeness. The workflow is substantially incomplete or incorrect, core code does not execute, or the implementation does not address the stated task. 30%
Reproducibility & Environment Management A new user can reproduce the workflow using clear setup and execution instructions, documented dependencies and versions, portable paths, explicit data-access steps, and controlled randomness where relevant. Restricted data is handled safely. The workflow is reproducible overall, with only minor gaps in dependency versions, setup instructions, paths, data access, or execution order. The main files are available, but missing dependencies, undocumented data steps, environment assumptions, stale outputs, or path problems make reproduction difficult. The repository lacks the files or instructions needed to recreate the environment, obtain the data, or run the workflow. 25%
Evaluation & Testing Evidence Baselines, splits or test cases, metrics, saved outputs, and error analysis are appropriate for the task and implemented correctly. Checks address leakage, invalid inputs, failure modes, and uncertainty where relevant. Evaluation is appropriate and implemented correctly overall, with minor omissions in comparisons, test coverage, examples, or failure analysis. Some evaluation evidence is included, but weak baselines, limited tests, questionable metrics, possible leakage, or missing error analysis reduces confidence in the results. Evaluation and testing are absent, incorrect, or insufficient to determine whether the implementation works as claimed. 25%
Code Quality, Documentation & Responsible Engineering The repository has a clear structure and entry point; code is readable, appropriately modular, and free of unnecessary duplication or exposed secrets. The README documents purpose, architecture, data and model provenance, limitations, responsible-use considerations, and AI assistance. Code and documentation are clear and professional overall, with minor issues in organization, naming, modularity, provenance, limitations, or AI-use documentation. The project can be interpreted with effort, but weak organization, monolithic or duplicated code, sparse comments, incomplete README guidance, or missing provenance and limitations reduces maintainability. The repository is disorganized or unsafe, documentation is substantially missing, secrets or restricted material are exposed, or the submitted code cannot be understood or responsibly reused. 20%