Task Definition & Intended Output
Specificity of the input, analytical operation, output, illustrative example, intended use, and initial definition of success.
Complete an end-to-end text-analysis project using a real-world collection such as news articles, social media posts, policy documents, or another appropriate source. The project should transform raw, unstructured text into analysis-ready data; implement a reproducible Python workflow for a course-aligned task; assess the quality and limitations of the results; and communicate the findings clearly through a poster and brief verbal explanation. Appropriate applications include classification, clustering, topic modeling, semantic search, retrieval-augmented generation, and question answering.
You may complete the project individually or in a group of up to four people. Group work is highly encouraged.
Begin by submitting a Topic Brief in one shared project repository. Continue using that same repository as you develop the full project. Use Python and, where appropriate, Google Colab and Hugging Face. Submit the Topic Brief repository URL, poster, and final GitHub repository URL through their designated Canvas links; work submitted by email will not be graded. The poster is due Wednesday, November 25, 2026, at 11:59 p.m. Late work is accepted with a 25% deduction for each day submitted late.
| Component | Percent of Final Grade |
|---|---|
| Topic Brief | 5% |
| Project poster presentation | 25% |
| Documented GitHub repository | 20% |
Work with your project team to create one shared GitHub repository and use its root README.md to introduce your proposed text-analysis project. The brief should make the task, intended output, preferred data, and immediate next steps understandable enough for the instructor to give useful feedback. This is an initial plan: you may revise the topic, data, or methods as the project develops.
One repository per team: Download the Markdown template, place it at the root of the team’s project repository as README.md, and submit that repository’s URL through Canvas. Continue using the same repository for the final project. The repository must be accessible to the instructor.
Zero-credit naming rule: The repository name must begin with the case-sensitive prefix DSA495, such as DSA495-news-classification. Check the name before submitting the URL.
The Topic Brief is graded for completion and specificity. Full credit requires the correct repository name, one accessible team repository with the completed template at its root, a concrete task with an illustrative input and desired output, a clear success statement, a sufficiently detailed and accessible preferred corpus, a feasible minimum project and immediate action plan, thoughtful questions or an explicit statement that the team has none, and a complete AI-use statement. The optional methods and evaluation section is not required for full credit.
These criteria total 100% of the Topic Brief component.
Specificity of the input, analytical operation, output, illustrative example, intended use, and initial definition of success.
Clarity of the corpus, original source and collection, verified access, unit of analysis, relevant fields, examples, scale, and data-use constraints.
Fit between the task and data, a realistic minimum project, recognition of consequential risks, and concrete assigned next actions.
Correct repository setup, required README sections, team identification, accessibility, source links, and AI-use documentation.
The performance band indicates the percentage of points earned within a criterion; the Weight column determines that criterion’s contribution to the component grade.
| Criterion | Exemplary 95–100% |
Proficient 80–94.9% |
Developing 40–79.9% |
Needs Improvement Below 40% |
Weight |
|---|---|---|---|---|---|
| Task Definition & Intended Output | The one-sentence task identifies a concrete input, operation, and output. The illustrative example makes the proposed behavior clear, and the intended result, use, and initial definition of success are precise, coherent, and technically plausible. | The task, illustrative example, intended output, and success statement are clear overall. Minor ambiguity remains about the input, operation, output format, intended use, or what would count as success. | The general topic is apparent, but the task is broad or the illustrative example, intended output, intended use, or success statement is incomplete or unclear. | The submission names a topic without defining an actionable text-analysis task or meaningful intended output. | 30% |
| Corpus Specification & Traceability | The corpus is identified with its original source and collection process, a direct link, a documented access check, and representative examples. Unit of analysis, relevant text and fields, approximate scale, task fit, and licensing or privacy constraints are documented accurately. | The corpus, provenance, access status, and relationship to the task are suitable and understandable, with only minor gaps in examples, collection details, fields, scale, or usage constraints. | A possible dataset is named, but its original source, collection, contents, access status, examples, or relationship to the task remains insufficiently verified. | No usable corpus is identified, the source cannot be verified, or the proposed data does not contain the information required by the task. | 35% |
| Scope & Feasibility | The task and corpus fit one another and support a realistic minimum viable project. The brief identifies the largest consequential risk, provides a workable backup plan, and assigns two concrete next actions with owners and approximate target dates. | The proposal appears feasible and appropriately scoped overall, with a plausible minimum project and action plan. A risk, backup plan, owner, or target date could be specified more clearly. | The idea may be workable, but the minimum project, major risk, backup plan, or immediate responsibilities are incomplete, leaving unresolved scope, access, labeling, computing, or evaluation concerns. | The proposed work is not feasible with the identified data, time, or course tools, and the brief does not identify a viable minimum version or executable next step. | 20% |
| Repository Readiness & Completeness | The accessible shared repository follows the required naming rule, places a complete README at the root, identifies all team members, links sources, and includes a complete AI-use statement. | The repository and README meet the core requirements, with only minor omissions or organization issues that do not impede review. | The repository is accessible, but several required sections, links, team details, or documentation elements are incomplete or difficult to locate. | The repository is inaccessible or substantially incomplete. A repository name that does not begin with the case-sensitive prefix DSA495 receives zero credit for the Topic Brief. |
15% |
Present the project in a poster format as a concise, audience-centered explanation of the analytical purpose, text data, technical workflow, principal findings, and practical implications. Use readable and accurate visualizations to make the text-analysis results accessible without obscuring important limitations.
Review the poster example provided for this course (opens in a new tab). For additional examples, see the bottom of this page (opens in a new tab).
Be prepared to explain the code, modeling decisions, evaluation strategy, and limitations of the project and to answer questions about how the submitted work was produced. Include the required AI-use statement in the submitted materials.
These criteria total 100% of the Poster Presentation component.
Precision of the task, data pipeline, representations, baselines, models, and important implementation decisions.
Quality of the comparison design, metrics, test evidence, error analysis, and attention to leakage, uncertainty, and limitations.
Accuracy, readability, visual hierarchy, purposeful figures, and the ability to communicate technical evidence concisely.
Ability to explain design decisions, code, results, failure modes, and individual or team contributions and to answer questions accurately.
The performance band indicates the percentage of points earned within a criterion; the Weight column determines that criterion’s contribution to the component grade.
| Criterion | Exemplary 95–100% |
Proficient 80–94.9% |
Developing 40–79.9% |
Needs Improvement Below 40% |
Weight |
|---|---|---|---|---|---|
| Problem & Technical Workflow | The poster defines the computational task precisely and presents a coherent end-to-end workflow. Data preparation, representations, baselines, models, and consequential implementation choices are technically sound and easy to trace. | The task and workflow are correct and understandable overall, with minor omissions in implementation detail, rationale, or connections between stages. | The main approach is visible, but important pipeline stages, baselines, model choices, or implementation decisions are incomplete, unclear, or technically questionable. | The task or workflow is substantially unclear or incorrect, and the audience cannot determine how the system produced its results. | 25% |
| Experimental Evidence & Evaluation Validity | The evaluation design matches the task and uses appropriate baselines, splits, metrics, examples, and error analysis. Claims follow from the evidence, and leakage, uncertainty, failure modes, and important limitations are handled rigorously. | Evaluation evidence is appropriate and supports the main claims, with only minor gaps in comparisons, metrics, error analysis, or discussion of validity threats. | Some results are reported, but weak baselines, limited test evidence, inappropriate metrics, possible leakage, or incomplete error analysis reduces confidence in the conclusions. | Evaluation is absent, invalid, or misleading, or claims about system quality are unsupported by reproducible evidence. | 30% |
| Poster Design & Technical Communication | The poster has a strong visual hierarchy and uses concise text, readable figures, accurate labels, and well-selected examples to communicate the technical work and results without oversimplification. | The poster is clear and readable overall. Visuals and explanations communicate the main technical story, with minor issues in density, labeling, organization, or emphasis. | The main message can be found, but crowded text, weak organization, default or unclear visuals, missing labels, or imprecise explanations make the technical evidence difficult to follow. | The poster is incomplete, inaccurate, or visually inaccessible and does not communicate the project’s task, workflow, evidence, and conclusions effectively. | 25% |
| Oral Defense & Team Command | The presenter or team explains design decisions, implementation, results, and limitations accurately and concisely. Responses to questions demonstrate direct command of the submitted code and clear individual contributions. | The presenter or team explains the project and answers questions accurately overall, with minor uncertainty about details or contributions. | Explanations show partial understanding, but important design decisions, results, code behavior, or team contributions cannot be explained clearly. | The presenter or team cannot explain essential parts of the submitted work, answer basic technical questions, or establish meaningful participation in the project. | 20% |
Submit a well-organized GitHub repository that contains the complete project workflow, from raw text through interpretable results. The repository may include multiple notebooks, scripts, configuration files, and supporting materials, but it must provide a clear entry point and enough documentation for another reader to understand and reproduce the work.
Include an AI-use statement in the repository. Follow the course AI policy, retain evidence of your process, protect confidential or restricted data, and be prepared to explain and reproduce your code, analytical choices, results, and any AI-assisted work.
These criteria total 100% of the Documented GitHub Repository component.
Correct implementation of the complete pipeline from source data through preparation, modeling or application, and final output.
Reliable setup instructions, pinned or documented dependencies, portable paths, data access procedures, and deterministic execution where appropriate.
Task-appropriate baselines, metrics, validation logic, saved outputs, error analysis, and checks that support claims about system behavior.
Readable modular code, clear repository structure, useful documentation, data and model provenance, limitations, and transparent AI use.
The performance band indicates the percentage of points earned within a criterion; the Weight column determines that criterion’s contribution to the component grade.
| Criterion | Exemplary 95–100% |
Proficient 80–94.9% |
Developing 40–79.9% |
Needs Improvement Below 40% |
Weight |
|---|---|---|---|---|---|
| Software Correctness & Complete Workflow | The repository implements a complete, technically sound pipeline from source data through final output. Core code executes successfully, algorithms are used correctly, and outputs are consistent with the stated task. | The complete workflow is present and correct overall. Minor implementation issues or omissions do not materially change the main results. | Major stages are present, but incomplete code, technical errors, hard-coded assumptions, or disconnected steps limit the workflow’s correctness or completeness. | The workflow is substantially incomplete or incorrect, core code does not execute, or the implementation does not address the stated task. | 30% |
| Reproducibility & Environment Management | A new user can reproduce the workflow using clear setup and execution instructions, documented dependencies and versions, portable paths, explicit data-access steps, and controlled randomness where relevant. Restricted data is handled safely. | The workflow is reproducible overall, with only minor gaps in dependency versions, setup instructions, paths, data access, or execution order. | The main files are available, but missing dependencies, undocumented data steps, environment assumptions, stale outputs, or path problems make reproduction difficult. | The repository lacks the files or instructions needed to recreate the environment, obtain the data, or run the workflow. | 25% |
| Evaluation & Testing Evidence | Baselines, splits or test cases, metrics, saved outputs, and error analysis are appropriate for the task and implemented correctly. Checks address leakage, invalid inputs, failure modes, and uncertainty where relevant. | Evaluation is appropriate and implemented correctly overall, with minor omissions in comparisons, test coverage, examples, or failure analysis. | Some evaluation evidence is included, but weak baselines, limited tests, questionable metrics, possible leakage, or missing error analysis reduces confidence in the results. | Evaluation and testing are absent, incorrect, or insufficient to determine whether the implementation works as claimed. | 25% |
| Code Quality, Documentation & Responsible Engineering | The repository has a clear structure and entry point; code is readable, appropriately modular, and free of unnecessary duplication or exposed secrets. The README documents purpose, architecture, data and model provenance, limitations, responsible-use considerations, and AI assistance. | Code and documentation are clear and professional overall, with minor issues in organization, naming, modularity, provenance, limitations, or AI-use documentation. | The project can be interpreted with effort, but weak organization, monolithic or duplicated code, sparse comments, incomplete README guidance, or missing provenance and limitations reduces maintainability. | The repository is disorganized or unsafe, documentation is substantially missing, secrets or restricted material are exposed, or the submitted code cannot be understood or responsibly reused. | 20% |