$ whoami
Headshot of Arsema Bekele

Arsema Bekele

I turn messy data into something you can trust. Information Systems student at Morgan State, and Data Architecture Extern at Breaking Games.

Open to internships
$ cat about.md

The Query Behind the Analyst

I'm Arsema, an Information Systems student at Morgan State University in Baltimore, graduating in May 2028. As a Data Architecture Extern at Breaking Games, I use SQL and Python to join, clean and validate messy data. What keeps me interested is the moment a confusing pile of records turns into something you can read, and the people and behavior behind it come into view. Outside of data, I'm into art and drawing and health and fitness, and I love getting things organized.

how_i_got_here.md

Data Lineage

I didn't grow up thinking I'd work with data. It started with trying to make sense of things in my own life: budgets, schedules, habits. I liked seeing the patterns.

Then I opened a dataset that looked impossible. Dates in three different formats, whole columns shifted out of place, random symbols where numbers should have been, duplicates stacked on duplicates. It wasn't pretty, but something about the mess pulled me in. I spent hours fixing it piece by piece, and when it finally snapped into something readable, it felt like watching static turn into a clear picture.

rawreadable

That's when I realized data isn't really about numbers. It's about understanding people, behavior and the stories hidden underneath the noise. It's the same curiosity that helped me separate bots from real customers in my Breaking Games externship: finding the truth inside the chaos.

eye_for_detail.md

Pattern Recognition IRL

I've been drawing since I was a kid, mostly portraits and quick sketches, all in pencil on paper. I like capturing faces, expressions and small details.

Drawing trains you to notice tiny things that feel off, and I think that's part of why I'm good at spotting patterns in data.

It also makes me think about layout and balance. When I design dashboards or build interfaces, I naturally care about spacing, clarity and how things feel visually.

I don't draw every day, but it's what I come back to when I need to slow down and reset.

Pencil sketch of a spotted big cat in profile, mouth open in a roar
$ git log --career

Career Commits

breaking_games.log

Data Architecture Extern

Breaking Games · Remote · July 2026 to present

orderscustomersunified customer view
  • Joined multi-source e-commerce datasets with SQL JOINs to build unified customer views
  • Investigated 1,000+ abandoned checkout sessions and separated bot activity from real customers, so the abandonment rate reflected real behavior
  • Standardized 421 referral-source strings into a channel taxonomy, and compared $37K in ad spend against $61K in sales to measure ROI
  • Modeled sales and inventory trends across 117 products to build holiday reorder recommendations
  • Wrote validation rules, process documentation and SOPs so the analysis can be repeated
SQLPythonExcelGoogle Sheets
tmcf_launchpad.log

TMCF Google Cloud Career Launchpad

11-week cohort · Compute Engine, Cloud Storage, IAM, BigQuery

Compute EngineCloud StorageIAMBigQuery
  • Weekly hands-on labs deploying, managing and securing cloud resources from Cloud Shell and Linux
  • Technical assignments, peer problem-solving, mock interviews and resume workshops
SELECT * FROM projects;

Build Log

nq_strategy_tester.py

Automated Strategy Testing System

Python · pandas · NumPy · Matplotlib · Databento API · Interactive Brokers API · In progress

trainheld-out test

Summary

I built a research pipeline for NQ futures, from raw market data through backtests and scoring to a paper-trading connection. I checked each strategy on data it was never tuned on, and two of three did not hold up. That is the point of the project.

What I built

  • A pipeline that pulled a year of 1-minute NQ data (350,503 rows) and flagged 3 degraded days
  • A backtester with commissions and slippage, plus equity curve, max drawdown and Sharpe metrics
  • A composite score that ranks strategies on risk and sample size, not just profit
  • A replay monitor that reproduced all 109 backtest alerts exactly, confirming no lookahead bias
  • A read-only Python connection to an Interactive Brokers paper account

Results (net NQ points, train → test)

  • Opening-range breakout, 5-min window: +292.65 → −144.10
  • VWAP mean reversion, 50-point threshold: +468.77 → +239.60
  • VWAP mean reversion, 70-point threshold: +498.92 → +103.95

The VWAP test sets are small (31 and 20 trades), so this is a promising lead, not proof. Backtests only, no live trading.

What I learned

  • A strategy that looks great in training can be noise. Both opening-range setups failed on unseen data.
  • A third split for tuning kept the final test clean. When a tuned stop-loss failed, the simpler version won.
  • Profit alone misleads. One MNQ run made $479 but had a $509 drawdown, so I treated it as not trade-ready.
doc_qa_chatbot.py

Document Q&A Chatbot

Python · Gradio · Google Colab

PDF

Summary

Upload a PDF, ask a question in plain English, and get the section that answers it.

Why I built it

Long documents are slow to search by hand, and I wanted to see how far simple keyword matching could get before reaching for anything heavier.

How it works

  1. Upload a PDF such as a certificate, packaging spec or contract
  2. Text is pulled out automatically
  3. Keywords are identified from what the user types in chat
  4. Sections of the document that contain those keywords are matched
  5. The most relevant snippet or a short summary comes back

Features

  • Extracts text from uploaded PDFs automatically, so documents can be searched in seconds
  • Picks out the keywords in a question and finds the matching sections
  • A dropdown lets the user choose how the answer comes back: Short Answer or Full Snippet
  • Built as a Gradio app with a chat panel, running in Google Colab
loan_quality_check.py

Bank Loan Data Quality Check

Python · Pandas · NumPy · Excel · Google Colab

rawcleaned

Summary

Audited and cleaned a 5,000-record, 14-column bank-loan dataset before analysis.

What I did

  • Audited the dataset for missing values, duplicate rows, inconsistent types, and formatting issues before analysis
  • Identified a credit-card spending field stored as text and designed preprocessing steps including numeric conversion, median/mode handling, whitespace cleanup, and categorical standardization
  • Exported a reusable cleaned dataset and documented each preprocessing step to support repeatable analysis on similar data
$ pip list

Skill Inventory

rawcleanvalidateinsight

Programming

Python, pandas, NumPy, Matplotlib, seaborn, SQL

Data

SQL JOINs and relational modeling, Excel (advanced formulas, pivot tables), Google Sheets, BigQuery, data cleaning and validation, pattern detection, visualization

Cloud

Google Cloud Platform, Compute Engine, Cloud Storage, IAM

Tools

Git, GitHub, VS Code, Jupyter, Linux, automated data pipelines, workflow documentation, quality control

Education

Morgan State University, B.S. Information Systems, Earl G. Graves School of Business. Expected May 2028. Database Management Systems, Data Structures, Information Systems Applications, Financial Accounting, Micro and Macroeconomics, Business Communications

Additional

Gradio, keyword search, process mapping, standard operating procedures

$ ls ~/offline

Outside the Terminal

> what is arsema listening to?

“Hold On”

by The Internet

1984GEORGE ORWELL
> what is arsema reading?

1984

by George Orwell

DrawingFitness
> what is arsema into?

Drawing and fitness

Drawing trains me to notice small details, and it's what I come back to when I need to slow down and reset. Fitness gives me structure: you show up, do the work, and feel better, and tracking small improvements keeps me grounded when life gets chaotic.