We will begin from the basics of the Python programming language and learn the key tools and libraries within Python needed to solve industry problems in data science. The course will include data manipulation and preprocessing using pandas and NumPy, cover machine learning models in scikit-learn, and conclude with building full preprocessing-and-forecasting pipelines. At the completion of this course, students will be able to write basic code and functions as well as visualize data, implement data science workflows, and fit machine learning models using Python. We will also explore differences between Python and R as tools for data science to help students select the correct tool for a given industry problem.Cannot be taken for credit after CS 161.
This foundational course offers a full-spectrum introduction to data science and data science workflows, emphasizing data as a source of value creation in the enterprise. The R programming environment serves as the implementation vehicle in support of essential data science activities - data exploration and visualization, data wrangling, predictive modeling, model deployment, and communication. The R programming environment, along with Python, is among the most important tools in the data scientist's toolbox and this course will feature tools and a style of programming inspired by the popular tidyverse ecosystem - ggplot2 for data visualization, dplyr and tidyr for data wrangling. Students will master elements of the data science workflow through a series of short R programming exercises reinforced by a full-spectrum, integrative final project. Presentation skills are an ever-present theme as students are challenged, through every stage of analysis, to communicate managerial relevance and value to the enterprise.
It's one thing to conduct an analysis, it's another to convince someone to change their behavior based on this analysis. In this course, students will study theories of visualization, communication and presentation with the purpose of translating technical results into actionable insight. Using a mix of case studies and code, the course begins with an examination of how to ask good research questions. It then covers the psychology of communication and the construction of compelling visualizations. Finally, students are tasked with writing and presenting their work in a manner suited to a non-technical audience.
Data management is core to both applied computer science and data science. This includes storing, managing, and processing datasets of varying sizes and types. This course introduces students to the various ways in which data is stored and processed including relational databases, file-based databases, cloud-based storage and data streaming. A key component of the course is learning which architectures fit which types of data science problem (and the strengths and weaknesses of each). Students will learn to work with data that is both clean and structured, and dirty and unstructured.
This course explores the legal, policy, and ethical implications of data. These types of issue arise at each stage of the data science workflow including data collection, storage, processing, analysis and use. Armed with legal and ethical guidelines, students are then confronted with topics including privacy, surveillance, security, classification, discrimination, decisional-autonomy, and duties to warn or act. Using case studies and a lecture-discussion format, the course will address real-world problems in areas like criminal justice, national security, health, marketing and politics.
Machine learning is becoming a core component of many modern organizational processes. It is a growing field at the intersection of computer science and statistics focused on finding patterns in data. Prominent applications include personalized recommendations, image processing and speech recognition. This course will focus on the application of existing machine learning libraries to practical problems faced by organizations. Through lectures, cases and programming projects, students will learn how to use machine learning to solve real world problems, run evaluations and interpret their results.
Over the course of the semester, students will propose, plan and execute an actual data science project. Run as an independent study during the student's last term, the project provides an opportunity to integrate all of the core skills learned throughout the program, and to develop a portfolio piece that can help with students' career aspirations. Projects must be consequential in nature-i.e., have a real (or potential) impact on some organization, or the world. Grades will be based on assessments by both the faculty advisor and those (potentially) impacted by the project's results. Data sets must be selected by the student either from a public repository or from the company for which they work and approved by the course instructor within the first two weeks of the term.
Data engineers design, implement, and maintain the data pipelines that power modern technology and society. This course builds on previous coursework, focusing on ingesting data from diverse sources, including logs, data streams, and various database types. Students will explore automated orchestration tools for scheduling and managing data pipelines, along with key concepts and technology in data warehousing and lakehouses. Emphasis will be placed on documentation, pipeline maintenance, and development through collaborative, semester-long projects.
This advanced machine learning course is designed to provide students an in-depth exploration of advanced topics, techniques, algorithms, and applications in the field of machine learning. Through a combination of lectures, hands-on exercises and project-based learning, students will gain a comprehensive understanding of machine learning techniques and their applications across domains. Topics may include: training and fine-tuning neural networks, generative networks, natural language processing, latent space representations, and evaluating ethical considerations such as dataset quality and bias. Particular emphasis will be given to current events and recent advances in the field.
We will begin from the basics of the Python programming language and learn the key tools and libraries within Python needed to solve industry problems in data science. The course will include data manipulation and preprocessing using Pandas and numpy, cover machine learning models in scikit-learn, and conclude with building full preprocessing-and-forecasting pipelines. At the completion of this course, students will be able to write basic code and functions as well as visualize data, implement data science workflows, and fit machine learning models using Python. We will also explore differences between Python and R as tools for data science to help students select the correct tool for a given industry problem.
In the first semester of this two-semester series, students will articulate their professional goals and aspirations. Through assessment and an exploration of their strengths, they will begin identifying how and where to focus their search for internships that utilize the skills and knowledge they're gaining through their Data Science curriculum. They will create application materials to market themselves, such as resumes, cover letters, and linkedin profiles, as well develop the core skills involved in effective networking. Additionally, use of AI during the job search process will be addressed. Students will develop lifelong career development skills and a plan that aligns with their skills, interests, and career goals.
In the second semester of this two-semester series, students will apply concepts and skills developed in the first semester, and begin engaging in active networking and researching organizations of interest. Students will practice skills for both traditional and technical interviewing, learn negotiating techniques, and take the final organizational steps towards launching a job or internship search. By the end of the course, students will have a personalized Career Action Plan and the skills and tools necessary to secure job opportunities and navigate their long-term career growth in the tech industry.
This experiential learning course is designed to support early-career students as they apply foundational data science knowledge in a real-world internship setting. Through a combination of hands-on professional experience and guided academic reflection, students will deepen their understanding of data science practices while developing critical workplace competencies. The course emphasizes career readiness, self-directed learning, and professional growth, encouraging students to identify, pursue, and refine personal and career development objectives using the SMART goals framework (Specific, Measurable, Achievable, Relevant, Time-bound). Students will engage in reflective assignments, peer discussions, and one-on-one coaching to make meaning of their experiences, strengthen their professional identities, and enhance their communication, collaboration, and problem-solving skills.
Survival analysis methods consist of statistical modeling techniques that predict the time to an event. The event of interest often depends on the application area. Survival analysis techniques have been used in a wide variety of areas-medical professionals working to predict the onset of a disease; HR representatives studying trends in workforce attrition; mechanical engineers investigating reliability of post-released products; and many other applications. In this course, students will be exposed to practice survival analysis techniques through exploration of real data sets. Students can expect to develop a deeper understanding of concepts including: observation censoring, accelerated life testing and experimental design, recurrence analysis, and implementation of survival analysis techniques in statistical software as well as communicating statistical results to different audiences.
Individualized program of investigative research, in which a student works directly with a Computer or Data Science faculty member on their area of research expertise. Nature of participation varies from collaborative research to the design and execution of an independent project. The course provides hands-on experience, which may include literature review, data collection, data management, data analysis, and the synthesis of results in a formal paper and/or oral presentation. May be repeated for credit until a maximum of 4 total credits.
A semester-long study of topics in Data Science. Topics and emphases will vary according to the instructor. This course may be repeated for credit with different topics. See the details in the schedule for descriptions and applicability to graduation requirements.
Willamette University
Ecotrust Building
721 NW 9th Ave
Portland
Oregon
97209
U.S.A.