Session 1B

Data Science, Statistics and Society

12:30 PM to 2:15 PM


Distributed Detection & Response to Mitigate Denial of Service Attacks
Presenter
  • Bikramjeet Singh Yashwinder (Bikram) Ghura, Senior, Information Technology (Tacoma)
Mentor
  • D.C. Grant, Institute of Technology (Tacoma Campus)
Session
  • 12:30 PM to 2:15 PM

Distributed Detection & Response to Mitigate Denial of Service Attacksclose

Distributed Denial of Service Attacks (DDoS) have been a prevalent threat to information security across organizations. Moreover, the rise of botnets has provided a lucrative environment for DoS attacks to evolve and increase at an alarming rate. Hence, it is crucial for organizations to integrate security policies and infrastructure within their operations. Various security products such as Intrusion Detection/Prevention systems (IDS/IPS) and firewalls help address this need. However, acquiring and configuring these products can be time-consuming and costly. Furthermore, configuring & integrating such systems together requires technical expertise in information security. This research study suggests a novel approach to developing a lightweight and low-cost sensor for early detection of DoS attacks. The sensor communicates over a distributed network, with an IPS system which responds by reducing the impact of DDoS attacks. Such a configuration also allows the solution to be scalable i.e. the IPS could be configured to manage data from multiple sensors. The project is being conducted in 3 phases: (1) Identify, implement and test the technologies required for monitoring and logging traffic to and from insecure device, within a virtual environment. (2) Extend configuration to the production environment including Raspberry Pi(s) acting as the sensors. (3) Add usability enhancements such as an interactive console or Graphical User Interface (GUI) for system configuration. With development in progress, it is anticipated that this project will open various avenues to explore the potential for effective use of remote devices to detect DoS attacks.


Multivariable Calculus Applications in Environmental Sciences
Presenters
  • Morgan Wolf, Freshman, Math, Physics , Lake Wash Tech Coll
  • Samuel (Sam) Wolf, Sophomore, Computer Science , Mathematics , Lake Wash Tech Coll
Mentor
  • Narayani Choudhury, Mathematics, Physics, Lake Washington Institute of Technology
Session
  • 12:30 PM to 2:15 PM

Multivariable Calculus Applications in Environmental Sciencesclose

Here we explore applications of multivariable calculus for studying three dimensional wave media in our environment. We employ regression based methods to derive analytic formulae for real wave. Using multivariable optimization methods, we derive the maxima, minima and saddle points of three dimensional functions. We use advanced data visualization methods to study the divergence and curl and illustrate how these can be used to study ocean waves- including their vorticity and circulation. The project provides hands on exploration of real world environmental science problems with advanced data visualization and shows how divergence and curl can be used to measure circulation and vorticity parameters of real wave media involving ocean waves. Real world manifestations of scalar and vector fields in our environment are also presented.


Monte Carlo Simulation Estimations of π
Presenter
  • Samuel (Sam) Wolf, Sophomore, Computer Science , Mathematics , Lake Wash Tech Coll
Mentor
  • Narayani Choudhury, Mathematics, Physics, Lake Washington Institute of Technology
Session
  • 12:30 PM to 2:15 PM

Monte Carlo Simulation Estimations of πclose

Monte Carlo simulations employ random probability distribution statistics to estimate areas and volumes. Here, we employ Monte Carlo simulations to estimate the numerical value of π. We inscribe a circle in a square board and throw N darts using random values for both x and y. The probability that the dart lies within the circle = area of circle/area of square. This relationship allows us to estimate π. We wrote EXCEL/JAVA code for this research. The accuracy of estimated π is improved as the number of darts N --> ∞. This research allows us to combine mathematics, computer programming and data visualization to estimate π. Important applications of Monte Carlo simulations to find areas and volumes of complex objects including rivers, landscapes and organisms which cannot be represented by analytic functions will be discussed.


Matching Heterogenous Datasets Using Seeded Classification
Presenter
  • Jainul Vaghasia, Sophomore, Statistics, Computer Science
Mentors
  • Martina Morris, Statistics
  • Ben Marwick, Anthropology
Session
  • 12:30 PM to 2:15 PM

Matching Heterogenous Datasets Using Seeded Classificationclose

For those interested in understanding more about fatal shootings of civilians by police, a natural place to start would be data on the events. Such data is, however, surprisingly hard to access. There are no official sources, but there are several crowd sourced data sets available online. These data sets each contain some common and some unique fields, and they are stored in very different formats. Hence, these data sets require preprocessing before the data can be used for other practical purposes like analysis and visualization. In this research, we focus on cleaning such heterogenous data sets using statistical methods with major focus on data matching. Data matching refers to matching data from different sources and recognizing the matching entities. It is an essential step in data cleaning that helps in resolving the conflicting information provided by the data sets---clerical errors, missing data, differently stored data, etc. Traditional statistical procedures and classification algorithms are based on supervised learning which requires training data. However, it is quite common that the training data is not available at hand and it might not be cost effective to prepare one from scratch. We explore a two step unsupervised learning algorithm that might achieve the same performance as the supervised counterparts. In the first step, it prepares a seed data set containing data points that are relatively extreme matches and non-matches based on discriminating fields. This prepared data set can be used as training data to train classical classification algorithms---k-nearest neighbors, Support Vector machines, Random Forests---which in turn can be used to classify the remaining data points. In the process, we examine missing data, geocode matching, and feature selection methods that increase classification accuracy by selecting the most discriminating fields. Finally, we examine the prospect of generalizing this approach to suit differently aligned data.


Fatal Encounters with Police: Improving Public Access to Exploratory Data Analytics
Presenter
  • Madeline Ann (Maddi) Cummins, Sophomore, Pre Engineering
Mentors
  • Martina Morris, Statistics
  • Ben Marwick, Anthropology
Session
  • 12:30 PM to 2:15 PM

Fatal Encounters with Police: Improving Public Access to Exploratory Data Analyticsclose

National interest in the fatal shooting of civilians by police is growing, driven in part by several high profile cases that were captured on video. Several news organizations, including the Washington Post and the Seattle Times, have done in-depth reporting on local or national trends. While the data these organizations used has been made public, the technical skills needed to access and analyze the data present a barrier to public use. This project seeks to develop browser-based software that will support simple access to the data, along with tools for Exploratory Data Analysis (EDA), that will make it easier for others to learn from these data. For this project we are working in the programming language R, writing a shiny app to provide the browser-based interface to R’s powerful EDA tools. Shiny apps involve writing two software components: the underlying R code for analyzing the data, and a user interface (UI) that provides a simple point and click tools for running the code. For the R code, we will focus on graphical exploration of the data: time series charts at different levels of spatial aggregation, interactive maps where users can zoom in and view online links for individual cases, choropleth maps to visualize racial disparities, and animated cartograms. For the UI, we will develop a webpage with tabs for each graphical option, and a range of filters and aggregation options on each tab. Users will be able to easily view and interact with each visual and modify it to display what they are interested in. Our codebase will follow the guidelines for reproducible research, and be uploaded to a public GitHub repository. This will support both public use and public development; users interested in contributing to the codebase can clone our repository, and submit suggested edits via pull requests.


The University of Washington is committed to providing access and accommodation in its services, programs, and activities. To make a request connected to a disability or health condition contact the Office of Undergraduate Research at undergradresearch@uw.edu or the Disability Services Office at least ten days in advance.