Cricut Design Space is the required software to use when interacting with a Cricut device. It is also one of the most infuriating, piece of shit softwares I have ever had the necessity to use. 
 Most of what makes it awful is a licensing decision rather than a technical one. A Cricut is a stepper-driven gantry with a tool head on it. That is a CNC machine. Nearly every other machine in that…
For most of the time I’ve been writing software, the thing that decided what actually shipped was how fast the team could build it. Engineering capacity — the throughput of design, build, test, and release — was the ceiling. It’s why engineering leadership tended to end up as the de facto gatekeepers of the roadmap: the constraint on what reached customers was the engineering…
If you’ve ever needed to work on multiple features simultaneously, you’ve likely encountered the friction of context switching between git branches. The traditional approach—stashing changes, switching branches, and checking out different code—disrupts your flow and can be error-prone. Git worktrees offer an elegant solution, and they become even more powerful when you’re working…
living in San Diego working on a number of new things
Will mostly talk about technical things, primary in the Kubernetes, engineering practices, and Reliability space, maybe with some opinions of mine thrown in the mix.
I do technical advising for early-stage startups, primarily in the B2B SaaS space, with a focus on backend systems/product architecture, reliability, observability, cloud…
I provide consulting services within software engineering and management, primarily focused on small to mid-sized engineering and product teams, with a focus on backend systems/product architecture, reliability, observability, cloud operations, security, and preparing teams to have systems in place to facilitate scale. If you’re looking for advice, you can book an initial conversation via my…
Preview# Preview is a SaaS application for generating and managing dynamic and ephemeral Preview environments. Production-like environments are created to track every Pull Request, or manually for demos, manual testing, or training.
MergeDeps# MergeDeps is a Github application for managing Pull Request dependencies and avoiding accidental premature merges.
Redirector# Redirector is a…
The PreviewHQ product is a hosted ephemeral Preview Environment service. It contains both an internal build and deployment service, and logs from these services need to be available to users.
The initial implementation of these log streams was based entirely on Kubernetes pod logs. On a request for logs, the application backend:
Queried for the name of the Kubernetes pod that ran that…
For a while, under any real load, one of my services was returning 502 errors at an ~0.5% rate.
This is a service that is run in Kubernetes, with an Nginx Ingress Controller, and a Flask python application being run with uWSGI.
I was expecting these errors to be fairly easy to identify because I generate Request IDs in nginx (added as an X-Request-Id HTTP header to both the downstream…
DigitalOcean is a player in the Infrastructure-as-a-Service (IaaS) market, competing with options such as Amazon Web Services (AWS), Google Cloud Platform (GCP), as well as other “low cost” providers such as Linode.
Amoung these services, DigitalOcean has their “Managed Database” offering. One of the downsides of this offering is the terribly low connection limits,…
Software leveling is a complicated debate. What is a “Junior” Engineer? At what point is someone “Senior”? Every company has a slightly different outlook on this, and this is just my opinion; it’s the mental model I use.
Junior Software Engineer# As a Junior Software Engineer, an individual would be primarily getting small-to-medium tasks delegated by their…
I while back I wrote about how I managed Kubeconfig files in my environment, in a easy-to-modify way.
I’ve changed my approach recently to make this easier to use; requiring a KUBECONFIG_MANUAL environment variable in order to avoid an explicitly-set KUBECONFIG being overwritten was error-prone and ergonomically poor.
I’ve recently decided to take a different approach, with the…
Container (or deployment) orchestration is the automation of taking your application, and placing it on infrastructure (typically a machine or VM) to execute as part of your larger application, while ensuring that the workload continues to run through a variety of potential issues. Scheduling, a subset of orchestration, determines which node (or nodes) in your infrastructure the application should…
As the third project in my 12 Startups in 12 Months year, I’m announcing Hookshot.
Hookshot is a managed platform for Webhook delivery. It is designed to handle the full lifecycle of a Webhook implementation, from allowing the end-user to create, manage, and inspect events, to ensuring reliable and secure delivery across a variety of potential issues.
Implementing a reliable webhook…
This is the first post in my new “Everyday Automation” series, which explores using free tools and services to make every day life just a little bit easier.
As a bit of background, through 2020 I have been fostering kittens. While fun and entertaining, the process of getting them actually adopted during COVID can be time consuming. The program we foster through switched to virtual…
Programming (or coding) is an extremely powerful skill that is becoming more important each year. There is a misconception though: that the benefits of these skills are always trapped behind learning a programming language, and writing your own software. Automation, one of the large benefits that programmers are able to utilize, can be accomplished in every day life using readily available tools.…
As the second project in my 12 Startups in 12 Months year, I’m releasing (sort of) RecruitMe.
I’m a full time software engineer, and week-to-week I’ll receive up ~5 cold-outreach opportunities from recruiters for new opportunities. Unfortunately the majority of these are clearly not a good fit, are impersonal, and are low-effort outreach from low-quality recruiters. If the…
“12 Startups in 12 Months” has been a relatively popular goal for a little bit now. The idea is to give yourself 12 months to attempt to start 12 unique companies, in a “fail fast” approach, with the goal of finding something that sticks.
This is, in my opinion, a good goal. It splits evenly into a year (1 concept per month), and gives you a set amount of time for each…
A core tenant of application promotion is that you are using the same version of your application between environments. If you have a version of your code running on staging, or in a Preview environment, and you want to promote that code to production, you simply deploy same image to production, with the production set of environment variables. With a populate-on-build system like Reacts’,…
Last month I released MergeDeps in order to allow developers and teams to easily blocked Pull Request merges based on dependencies.
Today, I’m announcing that MergeDeps can now block merges based on time.
Often times you may not want a PR to be merged before a specific date or time. Especially with a Continuous Deployment pipeline in place, you do not want to merge code into your…
Publishing of this post was postponed until the report vulnerability described here had been corrected. The core risk of software supply chain issues is still a widespread issue.
tldr; don’t depend on un-pinned external projects controlled by someone who doesn’t work at your company. Especially as a core part of your payment processing infrastructure.
Fun security-related thing that I…
The last couple days I’ve been involved in a couple unique conversations about the difference between Product and Project Management. As someone who does not work in either of these roles, but often works with people in these roles, these are just my opinions and expectations as an engineering looking in. I don’t expect to be 100% right, and recognize that these roles will have…
I hate doing this. Apparently writing public blog posts and crossing my fingers (🤞🏼) is the best way to get a delete Mailchimp account recovered.
This is a follow on from this post from 2018. And from this post in 2019. It sucks that I have to be the one writing the 2020 edition.
tldr; Mailchimp deleted my account and all of my data without notification because they felt I wasn’t…
There were decades of campaigns to get kids to wear helmets while riding bikes. Decades. 2018. 2017. 2015. 2010. 2000. 1990.
Then Bird/Lime/Uber/Lyft/whoever all came in and were like ELECTRIC SCOOTERS FOR EVERYONE NO HELMET NEEDED and just threw all of that progress in the trash?
I was out for a walk through downtown Austin, and I passed no-less-than 8 couldn’t-be-older-than-13-year-old…
Today, I am announcing the initial launch of a new project, MergeDeps. In one line, MergeDeps allows a developer to specify a dependency for a Github Pull Request, and MergeDeps will ensure that the Pull Requests are merged in the correct order. These dependencies are specified directly in the Pull Request body, without any necessary changes to your code or workflow.
Whether you are a solo…
First off, I apologize for the clickbait title. It hurt me just writing it.
A common step in a Continuous Integration/Continuous Delivery (CI/CD) pipeline is building container images. Fast image building is heavily dependent on being able to use a layer cache, which is many cases happens by default. The layer cache is what allows a docker build to skip a complex or long-running build step, by…
Large tech organizations are announcing permanent “work-from-home” policies. Companies are announcing that they will be “remote first”. As employees chose to move out of high cost of living areas, companies are introducing compensation adjustments. This is a very complex topic.
Overall, I feel that location-adjusted compensation for location-independent work is, at…
Disclaimer: I am currently building a product in this space at Preview. As such, I have biased opinions regarding the “correct” way of addressing Continuous Product Review needs. Preview is explicitly not included in this posting.
This article is meant to be a summary of existing tools in the space, and not a comparison between tools.
Over the last year, there as been an…
Over 2 years ago I wrote a quick Kubernetes controller in order to ensuring that “sidecar” containers were shut down after the “main” container in a pod exited. The issue we were seeing was fairly straight forward: we had a container running in a pod to accomplish some application logic, as well as a number of “sidecar” containers to provide some functionality…
Namespaces are a core component of the Kubernetes landscape which are often used as as a base level of resource isolation. As a resource which contains multiple others, the shutdown behaviour associated with terminating namespaces is complex. In effect, namespaces can often get “stuck” in a Terminating state. I ran into this recently with a variety of namespaces in a Digital Ocean…
Early today an engineer at Buffer put out a post about removing resource limits from a subset of their Kubernetes deployments in order to “make their services faster”. In summary, they were experiencing a kernel bug in which cpu throttling was being applied to containers which had not yet hit their CPU limits. They determined that this bug did not have an impact if no limits were set.…
Note that while Google took similar action, I will be focusing primarily on a discussion of Apple’s market dominance and anti-competitive behaviour, rather than Google’s. This is, honestly, because much of this post was written before Google took action.
There are many things at play here, and they are conflated. Early today Epic Games fired the first shots to kick of a war with…
On August 6th, 2020, Donald Trump signed an Executive Order prohibiting any transaction related to WeChat within the United States after September 15th, 2020. It was claimed to be in the name of “National Security”, however it feels more like a targetted attack to reduce day-to-day communication between the US and China.
This is a follow on from the Executive Order targeted at…
When I first tried to start writing more, I purposely attempted to avoid political topics. Those days are over.
I figured that the political topics were totally disjoint from the technical side of things that I was trying to talk about, and my opinions on political topics were not necessary. I don’t believe that assumption holds any more. Technology is being more and more a political…
In multiple situations and clusters I’ve encounted the problem of logs being lost from short-lived containers; ie, a container which is dynamically spun up to complete a single job, and then exiting. These containers will often only exist for a couple of seconds, which makes effectively collecting and forwarding their logs challenging.
There are two main log-collection methods in…
On May 15th, 2020, Facebook announced that it had reached an agreement to purchase Giphy, for $400M.
To many, this seems like a staggering amount of money for something that just serves GIFs. Facebook’s main business is in advertising, and GIFs are not inherently valuable for advertising purposes. There may be some possible value, such as Giphy selling “preferred placement”…
Note that I have no experience in the restaurant industry, and these are my thoughts as pure outsider. If you’re looking for an opinion with hard facts backing it up, you’re not going to find it here.
Over the past couple of days, I have been seeing a lot of discussion about how COVID-19 is the beginning of the end for the restaurant and bar industry. It is no secret that…
Personally, I believe that Instacart is attempting to operate on a fundamentally flawed business model, and is doomed to fail unless the underlying model is changed. It is being kept afloat by external money, but will never be economically viable in a socially responsible way. If they can’t make enough margins to actually pay their employees, it’s not a viable business.
Note that I have 0…
CircleCI is a great CI/CI solution for companies who do not want to build an internal expertise on running these systems. For small-to-medium startups (up to ~50 engineers), these is almost always a small tradeoff. Despite some issues I have brought up in the past, I still believe that CircleCI can be effectively used as long as it is setup and administered appropriately for the security concerns…
Note that in order to avoid confusion with this separate issue, the publish date of this post was pushed back.
On July 29th, 2019, I discovered what I consider to be a major security risk with the use of the CircleCI CI/CD product that should be unacceptable to any corporation or open source project. This article details how commonly used CircleCI features interact with GitHub security…
Many of us have been there. We interview at hot new companies, excited to join a small team building out a new product or platform. One of the major draws of joining a startup is being able to have a major contribution to the vision, and to have a level of transparency that you can’t get at a larger organization. You decide on one of these companies after hearing from everyone you spoke with…
Depending on your development workflow, rebasing may be a common component of your workflow; ie:
You branch off master (we’ll call this branch A) You add a bunch of commits, for instance, developing a new feature You branch off again for a small fix (we’ll call this branch B) B1---B2 branch-b / A1---A2---A3 branch-a / M master Now, A gets merged into master. Typically (if keeping a…
Note: There is an updated version of this functionality described in a later post.
If you deal with Kubernetes, you are most likely familar with the kubeconfig file, typically located at ~/.kube/config. This file typically contains the clusters, users, and contexts (combinations of clusters and users) that you use in order to connect to our Kubernetes environments (typically with…
In any engineering environment, linting is an important tool to help maintain style, and prevent simple mistakes. The same is true of your Helm charts. Chart linting is an easy tool that you can add to your pipeline to ensure your deployments are valid and versioned correctly.
Getting linting set up for your custom Helm charts is actually extremely straight-forward thanks to the Helm…
For personal projects, I run a single node “cluster” on an OVH node, with kubeadm. Over the last couple months, I’ve been having repeated issues with pod evictions, which are typically not an issue, except when affecting critical pods, such as the kube-apiserver and kube-scheduler. Despite these being marked as critical pods, they were still being evicted due to disk pressure.…
We recently set up PGAnalyze as a way of doing a quick database health check, and managed to get a couple quick wins out of it.
By looking at the Query Performance section, we were able to quickly identify a bad outlying query which was accounting for ~7.5% of our total DB activity. Furthermore, we were able to see that this query was loading ~7.65 GB of data per call, with the majority of the…
I’ve recently been making a switch from using Amazon’s Application Load Balancers (ALBs), to using their new Network Load Balancers (NLBs). There are many reasons for this switch (that are not the subject of this post), but there was one glaring short fall: We heavily rely on request log data centered around the ALB provided X-Amzn-Trace-Id Request ID value.
When switching to the…
Recently a small cluster I maintain became unresponsive due to a failure of the kube-apiserver. Kubeadm clusters are currently limited to a single master, this meant that any interaction with the cluster was impossible.
Due to interaction through the API being (obviously) impossible, troubleshooting and recovery required digging a bit deeper. By ssh -ing to the physical node, I was able to at…
Multi-tenancy systems are very common for a simple reason: it both saves costs, and reduces complexity. Unfortunately, the biggest downside is that in certain situations, a small group of clients can negatively impact the rest of the system.
At Cratejoy, we run with a multi-tenant set up which is part SaaS, part Marketplace, and part Website Hosting. These 3 sections of our business have…
Monitoring is a a large, and complex topic, so rightfully is a a lot written about it. One topic that is not covered often however, is effectively monitoring your database indexes. This particular aspect of monitoring is something that had not previously occurred to me as something important, however one incident was rather eye-opening.
Despite our best efforts, systems sometimes require downtime for a variety of reasons. Different systems that I’ve seen have built this capability into different areas of their stacks, mostly in the application itself, or into the web server (such as nginx). Many of these solution still require the application to be running in order to serve an appropriate maintence page. At Cratejoy, this…