arXiv Preprint · 2026

Exploring Vision-Language Models for
Open-Vocabulary Zero-Shot Action Segmentation

Asim Unmesh1, Kaki Ramesh2, Mayank Patel1, Rahul Jain1, Karthik Ramani1

1Purdue University, USA   ·   2Birla Institute of Technology and Science (BITS) Hyderabad, India

Prior temporal action segmentation methods train a model on a fixed dataset action vocabulary and evaluate on the same closed label set.

(a) Existing approaches — closed vocabulary, trained per dataset

OVTAS runs a frozen vision-language model over the video frames together with an inference-time action vocabulary and produces a segmentation without any training.

(b) OVTAS — open vocabulary, zero-shot, training-free

Problem setup. Existing methods (a) are trained against a fixed dataset action vocabulary \(\mathcal{A}_\mathcal{D}\) and do not generalize beyond it. OVTAS (b) takes an inference-time action vocabulary \(\mathcal{A}_{\text{inference}}\) and a frozen VLM \(\mathcal{M}_{\text{vlm}}\), and segments the video with no task-specific training or fine-tuning.

Abstract

Temporal Action Segmentation (TAS) requires dividing videos into action segments, yet the vast space of activities and alternative breakdowns makes collecting comprehensive datasets infeasible. Existing methods remain limited to closed vocabularies and fixed label sets. In this work, we explore the largely unexplored problem of Open-Vocabulary Zero-Shot Temporal Action Segmentation (OVTAS) by leveraging the strong zero-shot capabilities of Vision–Language Models (VLMs). We introduce a training-free pipeline that follows a segmentation-by-classification design: (i) Frame–Action Embedding Similarity (FAES) matches video frames to candidate action labels, and (ii) Similarity-Matrix Temporal Segmentation (SMTS) enforces temporal consistency. Beyond proposing OVTAS, we present a systematic study across 14 diverse VLMs, providing the first broad analysis of their suitability for open-vocabulary action segmentation. Experiments on standard benchmarks show that OVTAS achieves strong results without task-specific supervision, underscoring the potential of VLMs for structured temporal understanding. We release code and embeddings at our project page. [Features available now; code release in preparation.]

14
VLMs studied
SigLIP · CLIP · OpenCLIP · PECore
3
Benchmarks
GTEA · 50 Salads · Breakfast
0
Training steps
No backpropagation, no fine-tuning

Downloads

VLM features — available now

We release the pre-extracted per-frame VLM features for all 14 VLMs across all three datasets (GTEA, 50 Salads and Breakfast), together with the cached action-label text embeddings. These are the exact features behind every number reported on this page.

💾 Download VLM features action_seg_vlm_feats.zip · ~99 GB · Google Drive

Extracting features from large VLMs demands substantial compute — ours were extracted on an NVIDIA A6000. By releasing ready-to-use embeddings we hope to remove that barrier, so that the community can iterate on open-vocabulary action segmentation and other action understanding tasks without repeating the extraction cost.

The archive is large. Google Drive cannot virus-scan a file this size, so the browser will warn before downloading; for a transfer this big we recommend a resumable command-line client such as gdown or rclone rather than the browser.

Code — coming soon

The OVTAS codebase will be released publicly on this page. We are preparing the release now and will post the link here as soon as it is ready. It will include feature extraction, the FAES stage, the optimal-transport SMTS decoder, all four training-free baselines (Random‑Uniform, ES‑Mean, ES‑Vote, ES‑NRP), and the evaluation scripts for F1@{10,25,50}, Edit and frame accuracy.

Please check back here, or watch the arXiv listing for updates.

Method

OVTAS is a training-free, zero-shot pipeline that requires only a set of candidate action labels (the action-set supervision) and the video frames as input. Action-set supervision assumes that, given the high-level activity (e.g. “making tea”), we know the set of possible fine-grained actions (e.g. “boil water”, “pour tea”, “add sugar”) — but neither their order nor their boundaries.

The two-stage OVTAS pipeline. Stage 1 encodes video frames and action label text with frozen encoders and forms a similarity matrix. Stage 2 decodes that matrix into temporally consistent segments with optimal transport.
The OVTAS pipeline. A two-stage “segmentation by classification” design. Stage 1 (FAES) generates a similarity matrix by matching frames with action labels. Stage 2 (SMTS) uses optimal transport with a temporal prior to enforce temporal consistency, producing stable action segments. Both encoders remain frozen (❄).
Stage 1

Frame–Action Embedding Similarity

Each action label is normalised into a natural-language phrase (“pour_coffee” → “pour coffee”) and encoded by the VLM text encoder; frames are encoded by the vision encoder. Both are \(\ell_2\)-normalised row-wise and matched by cosine similarity.

Stage 2

Similarity-Matrix Temporal Segmentation

Frame-level VLM predictions are temporally inconsistent because they are made independently. SMTS decodes the similarity matrix into a temporally consistent label sequence with an entropy-regularised optimal transport formulation carrying a temporal prior.

Stage 1 — FAES

Let \(T\) be the number of frames, \(N\) the number of action labels and \(C\) the embedding dimension. With row-wise \(\ell_2\)-normalised frame embeddings \(\mathbf{X}\in\mathbb{R}^{T\times C}\) and action-label text embeddings \(\mathbf{A}\in\mathbb{R}^{N\times C}\), the similarity matrix and the per-frame classification probabilities are

\[ \mathbf{S} \;=\; \mathbf{X}\,\mathbf{A}^{\top} \in \mathbb{R}^{T\times N}, \qquad \mathbf{P} \;=\; \mathrm{softmax}_{N}(\mathbf{S}). \]

Stage 2 — SMTS

We adopt the Action Segmentation OT (ASOT) decoder of Xu and Gould. Given \(\mathbf{S}\), we define a visual cost \(\mathbf{C}=\mathbf{1}-\mathbf{S}\) and a diagonal temporal prior \(R_{ij}=\left|\tfrac{i}{T}-\tfrac{j}{N}\right|\), which encourages monotone alignment. We then solve for a coupling \(\Pi\in\mathbb{R}^{T\times N}\):

\[ \Pi^{\star}=\arg\min_{\Pi\in U(\mathbf{u},\mathbf{v})} \;\langle \Pi,\;\mathbf{C}+\rho R\rangle \;-\;\varepsilon\,H(\Pi), \qquad \mathbf{u}=\tfrac{1}{T}\mathbf{1}_{T},\;\;\mathbf{v}=\tfrac{1}{N}\mathbf{1}_{N}, \]

where \(U(\mathbf{u},\mathbf{v})\) is the transport polytope and \(H(\Pi)\) the entropy of \(\Pi\). The entropy term makes the problem convex with a unique solution; \(\Pi^{\star}\) is computed with log-stabilised Sinkhorn iterations, which scale linearly in \(T\!\times\!N\) per iteration. Each frame is finally assigned the action with maximum transported mass, \(\hat{y}_t=\arg\max_{j}\Pi^{\star}_{t,j}\).

What “training-free” means here

Training-free refers strictly to the absence of backpropagation-based weight updates. Both encoders are used off-the-shelf and are never fine-tuned. Optimal-transport hyperparameters are selected by grid search on a small set of held-out videos per dataset and then fixed for all experiments — we use \(\varepsilon = 0.07\), \(\alpha = 0.5\), \(r = 0.04\), \(\lambda_{\text{frames}} = 0.11\) and \(\lambda_{\text{actions}} = 0.01\). Unlike the unbalanced formulation of prior work, we find the balanced formulation gives the strongest results. Because the action order is unknown under action-set supervision, the ordering of action embeddings used to build the temporal prior is randomised.

Positioning against prior supervision regimes

Method Zero-ShotOpen-VocabUnsupervised Weakly Sup.Semi-Sup.Training-Free
ASAL✗✗✓✗✗✗
COIN-SSL✗✗✗✗✓✗
NN-Viterbi✗✗✗✓✗✗
ASAL-SSL✗✗✗✗✓✗
TCCNet✗✗✗✓✗✗
UDE✗✗✓✗✗✗
CTC✗✗✗✓✗✗
D3TW✗✗✗✓✗✗
ASRF-unsup.✗✗✓✗✗✗
C2F-TCN✗✗✓✗✗✗
HTK✗✗✗✓✗✗
BCN✗✗✗✓✗✗
HVQ✗✗✓✗✗✗
U-OT✗✗✓✗✗✗
OVTAS (ours)✓✓✗✓✗✓

Scroll the table horizontally →

Table 1. Comparison with other TAS methods by generalization ability and training-data needs. ✓ denotes applicability. OVTAS occupies a regime no prior method does: training-free, zero-shot and open-vocabulary at once.

Results

We evaluate on three standard action segmentation benchmarks using their official evaluation splits, reporting the average across splits. We use five metrics: F1@10, F1@25, F1@50, Edit and frame Accuracy; Avg is the mean of those five. All four baselines are themselves training-free, zero-shot and open-vocabulary, and share the same frozen encoder features and prompts as OVTAS.

Takeaway. OVTAS outperforms every training-free baseline by a wide margin on all three benchmarks — roughly 6.8× the best baseline Avg on 50 Salads, 3.3× on GTEA and 2.3× on Breakfast — establishing encouraging results for the novel task of open-vocabulary zero-shot action segmentation.
Method Backbone (#params) F1@10F1@25F1@50 EditAccAvg
Training-free zero-shot baselines
Random—0.000.000.000.205.461.13
ES-Mean—1.490.130.045.405.632.54
ES-Vote—2.770.840.316.725.683.26
ES-NRP—6.815.902.738.007.246.14
OVTAS (ours) — SigLIP
SigLIP-M1so400m-p16-256-i18n
877.96 M
42.631.414.188.731.541.7
SigLIP-M2large-p16-384
652.48 M
42.431.314.287.931.441.4
SigLIP-M3so400m-p14-384
1128.76 M
42.730.514.387.331.341.2
OVTAS (ours) — OpenCLIP
OpenCLIP-M1ViT-B/32 (laion2B-s34B-b79K)
151.28 M
40.830.013.278.030.738.5
OpenCLIP-M2ViT-B/16 (laion2B-s34B-b88K)
149.62 M
41.131.113.879.231.439.3
OpenCLIP-M3ViT-L/14 (laion2B-s32B-b82K)
427.62 M
39.229.212.476.030.437.4
OpenCLIP-M4ViT-H/14 (laion2B-s32B-b79K)
986.11 M
39.629.213.577.729.537.9
OpenCLIP-M5ViT-g/14 (laion2B-s34B-b88K)
1366.68 M
39.128.812.473.730.136.8
OVTAS (ours) — CLIP
CLIP-M1ViT-B/32
151.28 M
41.831.213.888.631.341.4
CLIP-M2ViT-B/16
149.62 M
42.630.914.588.031.541.5
CLIP-M3ViT-L/14
427.62 M
41.330.013.685.530.840.2
OVTAS (ours) — PECore
PECore-M1B/16-224
447.66 M
39.828.413.679.130.338.2
PECore-M2L/14-336
671.14 M
38.826.711.668.828.534.9
PECore-M3G/14-448
2419.27 M
39.328.513.375.229.337.1

Scroll the table horizontally →

Method Backbone (#params) F1@10F1@25F1@50 EditAccAvg
Training-free zero-shot baselines
Random—0.120.010.012.065.611.56
ES-Mean—7.083.991.6713.005.996.35
ES-Vote—7.324.631.8712.685.966.49
ES-NRP—7.755.352.018.2912.967.27
OVTAS (ours) — SigLIP
SigLIP-M1so400m-p16-256-i18n
877.96 M
21.113.04.951.228.323.7
SigLIP-M2large-p16-384
652.48 M
21.313.85.451.128.123.9
SigLIP-M3so400m-p14-384
1128.76 M
20.813.04.950.727.923.5
OVTAS (ours) — OpenCLIP
OpenCLIP-M1ViT-B/32 (laion2B-s34B-b79K)
151.28 M
15.29.63.634.020.416.6
OpenCLIP-M2ViT-B/16 (laion2B-s34B-b88K)
149.62 M
15.910.03.835.321.017.2
OpenCLIP-M3ViT-L/14 (laion2B-s32B-b82K)
427.62 M
14.59.33.432.720.416.1
OpenCLIP-M4ViT-H/14 (laion2B-s32B-b79K)
986.11 M
14.89.43.533.120.216.2
OpenCLIP-M5ViT-g/14 (laion2B-s34B-b88K)
1366.68 M
14.49.13.331.919.915.7
OVTAS (ours) — CLIP
CLIP-M1ViT-B/32
151.28 M
19.713.04.944.725.621.6
CLIP-M2ViT-B/16
149.62 M
20.113.35.145.225.921.9
CLIP-M3ViT-L/14
427.62 M
19.212.74.743.925.221.2
OVTAS (ours) — PECore
PECore-M1B/16-224
447.66 M
15.310.33.537.022.517.7
PECore-M2L/14-336
671.14 M
13.68.52.930.419.615.0
PECore-M3G/14-448
2419.27 M
14.29.13.231.820.115.7

Scroll the table horizontally →

Method Backbone (#params) F1@10F1@25F1@50 EditAccAvg
Training-free zero-shot baselines
Random—0.000.000.000.7217.453.63
ES-Mean—20.5611.973.6322.3917.1415.14
ES-Vote—22.1513.694.5224.6617.1516.43
ES-NRP—24.3819.468.7336.3511.8320.15
OVTAS (ours) — SigLIP
SigLIP-M1so400m-p16-256-i18n
877.96 M
54.039.515.092.730.946.4
SigLIP-M2large-p16-384
652.48 M
54.039.515.292.730.946.5
SigLIP-M3so400m-p14-384
1128.76 M
53.839.215.192.530.846.3
OVTAS (ours) — OpenCLIP
OpenCLIP-M1ViT-B/32 (laion2B-s34B-b79K)
151.28 M
52.638.214.890.330.445.3
OpenCLIP-M2ViT-B/16 (laion2B-s34B-b88K)
149.62 M
53.538.915.491.230.745.9
OpenCLIP-M3ViT-L/14 (laion2B-s32B-b82K)
427.62 M
52.838.114.990.530.245.3
OpenCLIP-M4ViT-H/14 (laion2B-s32B-b79K)
986.11 M
53.038.415.090.830.345.5
OpenCLIP-M5ViT-g/14 (laion2B-s34B-b88K)
1366.68 M
52.337.714.789.829.944.8
OVTAS (ours) — CLIP
CLIP-M1ViT-B/32
151.28 M
54.039.515.292.630.946.4
CLIP-M2ViT-B/16
149.62 M
54.139.615.392.831.046.6
CLIP-M3ViT-L/14
427.62 M
53.739.115.092.130.746.1
OVTAS (ours) — PECore
PECore-M1B/16-224
447.66 M
53.939.415.491.630.746.2
PECore-M2L/14-336
671.14 M
52.137.314.588.729.644.4
PECore-M3G/14-448
2419.27 M
53.038.515.189.930.145.3

Scroll the table horizontally →

Table 2. Comparison of baselines and all 14 VLM variants across the three datasets. Avg is the mean of the F1 scores, Edit and Accuracy. The highest Avg within each family is highlighted. Baselines are reported to two decimals as in the paper; OVTAS rows to one.

What the numbers say

  • Among the training-free baselines, ES-NRP is consistently strongest, confirming that equal-splits baselines benefit from a non-repetition constraint — so it is the most representative baseline to beat.
  • Differences between VLMs are small on Breakfast and 50 Salads but pronounced on GTEA, where the egocentric viewpoint, rapid transitions and large number of fine-grained classes make the benchmark hardest.
  • Edit scores sit far above the F1 scores throughout. Since Edit evaluates temporal ordering and continuity while F1 measures segment-wise overlap at a threshold, this gap suggests ordering is recovered more reliably than exact boundary placement.

Benchmark statistics

DatasetMin (s)Max (s)Mean (s)
GTEA42.27133.9374.35
Breakfast12.27649.53137.40
50 Salads251.60605.00385.25

Table 3. Video duration statistics, in seconds.

DatasetMinMaxMean
GTEA244936.3
Breakfast1185.4
50 Salads152720.7

Table 4. Ground-truth segments per video.

DatasetMin (s)Max (s)Mean (s)
GTEA0.0744.731.94
Breakfast0.07386.0020.95
50 Salads0.03138.4718.59

Table 5. Segment duration statistics. GTEA's mean segment lasts just 1.94 s against 20.95 s in Breakfast and 18.59 s in 50 Salads — the model must repeatedly localise boundaries within very short spans, leaving little room for temporal context aggregation.

Ablations

Stage ablation

We ablate Stage 1 by randomly permuting the frame and action-embedding features, and Stage 2 by taking per-frame argmax predictions instead of the OT decoding. Both stages prove critical: removing either produces significant drops across every metric on every dataset.

DatasetAblated stage F1@10F1@25F1@50EditAccAvg
50 SaladsNone (full pipeline) 41.8431.1913.8488.5831.3041.81
w/o Stage 1 (FAES) 6.71 ↓35.13 5.16 ↓26.03 2.27 ↓11.57 9.35 ↓79.23 28.53 ↓2.77 10.40 ↓31.41
w/o Stage 2 (SMTS) 0.17 ↓41.67 0.08 ↓31.11 0.03 ↓13.81 0.88 ↓87.70 5.48 ↓25.82 3.01 ↓38.80
GTEANone (full pipeline) 21.3313.765.4451.0728.0823.14
w/o Stage 1 (FAES) 1.38 ↓19.95 0.60 ↓13.16 0.25 ↓5.19 17.39 ↓33.68 17.30 ↓10.78 7.38 ↓15.76
w/o Stage 2 (SMTS) 1.18 ↓20.15 0.52 ↓13.24 0.16 ↓5.28 3.95 ↓47.12 5.41 ↓22.67 2.64 ↓20.50
BreakfastNone (full pipeline) 54.2639.6015.3492.4430.8946.51
w/o Stage 1 (FAES) 12.58 ↓41.68 8.36 ↓31.24 2.85 ↓12.49 18.87 ↓73.57 26.08 ↓4.81 13.75 ↓32.76
w/o Stage 2 (SMTS) 1.16 ↓53.10 0.67 ↓38.93 0.27 ↓15.07 8.13 ↓84.31 17.50 ↓13.39 5.95 ↓40.56

Scroll the table horizontally →

Table 6. Stage ablations. Ablating Stage 1 corresponds to Random-ASOT; ablating Stage 2 corresponds to per-frame argmax decoding. Drops from the full pipeline are shown in blue.

Temporal prior and \(\ell_2\) normalisation

We independently ablate the \(\ell_2\) normalisation in Stage 1 and the temporal prior in Stage 2, reporting on GTEA. Both design choices turn out to be load-bearing: without either one, the Avg falls to 2.69 and 3.95 respectively — close to the random baseline (1.56) and well below the strongest training-free baseline (7.27).

MetricBest w/o Temporal Priorw/o \(\ell_2\) Norm
F1@1021.33 1.12 ↓20.21 2.36 ↓18.97
F1@2513.76 0.56 ↓13.20 1.12 ↓12.64
F1@505.44 0.19 ↓5.25 0.62 ↓4.82
Edit51.07 5.04 ↓46.03 7.23 ↓43.84
Acc28.08 6.54 ↓21.54 8.43 ↓19.65
Avg23.14 2.69 ↓20.45 3.95 ↓19.19

Table 7. Ablation on the temporal prior and \(\ell_2\) normalisation, on GTEA. Arrows show the drop from Best.

Analysis

Which VLM family works best?

Averaging each family's performance across the three benchmarks yields a stable ranking: SigLIP outperforms all others, followed by CLIP, with OpenCLIP and PECore trailing. This ordering is consistent across all three datasets, suggesting the relative strength of each family is stable across domains.

Line plot of the average metric per VLM family across GTEA, 50 Salads and Breakfast, showing SigLIP highest, then CLIP, then OpenCLIP and PECore.
Figure 3. VLM family average of the Avg metric across datasets (GTEA, 50 Salads, Breakfast). Each line is a VLM family.

Does scaling the VLM help?

No. In every family the largest checkpoint is never the best one. OpenCLIP ViT-B/16 (149.62 M) beats ViT-g/14 (1366.68 M); CLIP ViT-B/16 beats ViT-L/14; PECore B/16-224 beats both L/14-336 and G/14-448; SigLIP so400m-p14-384 (1128.76 M) trails its two smaller siblings on all three datasets. Simply scaling VLMs up does not yield better action-segmentation performance under OVTAS.
Scatter of performance against model parameter count, grouped by family, with shaded parameter-size bins; larger models do not perform better.
Figure 4. Performance vs. model size. Models are grouped by family (SigLIP, CLIP, OpenCLIP, PECore). Shaded regions indicate parameter-size bins: Low (≤400 M), Large (400–800 M), Huge (800–1500 M) and Giant (>1500 M).

As future directions to close this gap, we point to (i) stronger text prompting and (ii) video-frame pre-processing such as cropping — together these may extract frame and action-label embeddings that do improve with increasing VLM size.

Where does OVTAS get confused?

Examining per-action confusion on GTEA (SigLIP-M1) shows that take acts as a dominant attractor for most manipulation actions, since its reaching motion overlaps with the onset of nearly every other action. shake is almost entirely misclassified, reflecting the limitation of image-trained VLMs on dynamic motions.

GT action type takeputopenclosepourscoop spreadstirfoldshakeSIL
Accuracy (%) 413635353721 415437122
Top misclassified as puttaketaketaketaketake putopenSILopenput
… (%) 121721291528 2915243129

Scroll the table horizontally →

Table 8. Action-type confusion on GTEA (SigLIP-M1).

Effect of video length and segment density

Performance decreases consistently as video length increases, across all three datasets: longer videos bring greater temporal variability and amplify error propagation in training-free segmentation.

Video lengthAccEditF1@10F1@25F1@50Avg
0–60 s46.3198.5176.0562.3929.1662.48
60–120 s40.9996.2863.9850.8920.2654.48
≥120 s26.4182.3741.4025.818.1636.83
Video lengthAccEditF1@10F1@25F1@50Avg
0–60 s27.7550.6524.3817.148.3825.66
60–120 s23.3133.1913.548.123.0516.24
≥120 s16.8212.996.252.270.577.78
Video lengthAccEditF1@10F1@25F1@50Avg
240–360 s30.5784.3742.1531.3913.7340.44
360–480 s32.9275.0141.7331.3414.9639.19
≥480 s23.4076.1128.3119.887.4131.02

Table 9. Performance variation by video duration.

The density of fine-grained action boundaries matters just as much. GTEA videos comprise many brief segments (mean ≈ 36 per video) and the model performs worst there; Breakfast averages ≈ 5 segments and performs best; 50 Salads sits between the two. Tightly packed sequences of short actions are especially difficult.

GT segmentsAccEditF1@10F1@25F1@50Avg
20–2928.2372.7748.5533.5316.1839.85
30–3923.1242.7919.1111.784.7820.32
40–4922.1915.1410.065.952.2711.12
GT segmentsAccEditF1@10F1@25F1@50Avg
0–442.7998.6971.9059.2826.2559.78
5–929.0088.5650.5235.0612.6043.15
10–1425.1172.6138.8123.167.8233.50
15–1922.1432.3021.5412.734.3518.61
GT segmentsAccEditF1@10F1@25F1@50Avg
15–1925.7383.7732.7923.8610.0635.24
20–2432.7477.4844.3733.1214.0840.36
25–2932.0065.5738.5724.6214.1834.99

Table 10. Performance variation by the number of ground-truth action segments per video.

Qualitative Results

Each pair of colour bars shows the ground-truth segmentation (top) above the OVTAS prediction (bottom) for one video, with segments coloured by action label.

GTEA qualitative result: ground truth and predicted segmentation bars.
GTEA qualitative result: ground truth and predicted segmentation bars.
GTEA qualitative result: ground truth and predicted segmentation bars.
GTEA qualitative result: ground truth and predicted segmentation bars.
50 Salads qualitative result: ground truth and predicted segmentation bars.
50 Salads qualitative result: ground truth and predicted segmentation bars.
50 Salads qualitative result: ground truth and predicted segmentation bars.
50 Salads qualitative result: ground truth and predicted segmentation bars.
Breakfast qualitative result: ground truth and predicted segmentation bars.
Breakfast qualitative result: ground truth and predicted segmentation bars.
Breakfast qualitative result: ground truth and predicted segmentation bars.
Breakfast qualitative result: ground truth and predicted segmentation bars.

Figure 5. Qualitative segmentation results of OVTAS. Top bar: ground truth. Bottom bar: prediction.

Limitations & Future Work

BibTeX

@article{unmesh2026ovtas,
  title   = {Exploring Vision-Language Models for Open-Vocabulary
             Zero-Shot Action Segmentation},
  author  = {Unmesh, Asim and Ramesh, Kaki and Patel, Mayank and
             Jain, Rahul and Ramani, Karthik},
  journal = {arXiv preprint arXiv:2602.21406},
  year    = {2026}
}

Acknowledgements

We acknowledge the Feddersen Distinguished Professorship Funds. This work was also supported by NSF under the Partnership for Innovation: Technology Transfer (PFI-TT) 2329804. Any opinions, findings, and conclusions expressed in this material are those of the authors and do not necessarily reflect the views of the funding agency.